0:04 Chris Anderson: Great to see you, welcome to TED.
0:06 Now, Silvana, you've been passionate about science for quite a long time.
0:10 Tell me about this picture.
0:13 Silvana Konermann: So this picture is when I was 15.
0:16 So I was born in a small town in Switzerland.
0:19 My parents weren't into science, but somehow I got really fascinated just with nature around me and also just how we worked as humans and our biology.
0:27 So I really wanted to find a way to be able to get into a lab to do some science.
0:32 It was actually pretty tricky for me, but eventually I talked one of my science teachers into convincing one of his colleagues to let me go into the lab.
0:40 And so this is me with that first science project where I went on to win the national competition, and then also the European Union competition.
0:50 And I think that's really where, you know, I got, I think, the confidence to continue with science since then.
0:57 CA: But there's a drawback to being a scientific prodigy, which is that you end up feeling like you might have a responsibility to do something with that.
1:04 And I think you've had that your whole life and you’ve thought about what are the biggest problems you could work on.
1:11 Tell us about this graph here.
1:13 SK: Yeah, absolutely.
1:15 So I have been, you know, doing science now for more than 20 years.
1:18 But I did want to say, this is actually my first real public appearance.
1:22 So I am very much, usually behind the scenes.
1:27 (Applause) Yeah, a problem that I’ve really been thinking a lot about since undergrad -- I did my undergrad in Switzerland in biology neuroscience -- and I learned more about Alzheimer’s disease.
1:41 And through learning how there are these big changes in the brain that are happening, a lot of what is known about late stages of the disease, how severe it is, and then the lecture though ended with basically: but we have no idea really how it’s starting.
2:01 We still don’t have a therapy.
2:02 And that was now, a long time ago, I guess, 17 years ago or so.
2:08 And that really stuck with me because why didn’t we understand how it’s starting?
2:13 Why didn’t we have a therapy?
2:15 There are all these very observable changes happening.
2:19 And so that got me interested in disease biology and specifically complex diseases, where Alzheimer’s is a complex disease.
2:27 It sounds “Oh, it’s just complicated,” but that's not what it means.
2:31 It means that there are multiple different risk factors.
2:34 And basically every patient has a unique combination of risk factors for a disease -- that’s different from an infection where you have one cause.
2:43 CA: And several of these diseases here are similarly fundamentally complex.
2:47 SK: That's right.
2:48 So heart disease, many cancers, obviously not accidents, but stroke and Alzheimer’s disease -- these are all complex diseases.
2:56 CA: And so that's why it's been so resistant to dramatic advancements in medical science in recent years.
3:04 SK: Basically, all of these have a combination of genetic changes environmental factors.
3:09 And each patient is unique.
3:11 They have a unique combination of risk factors.
3:13 And so we’ve been really struggling as a scientific community understanding what do all these different patients have in common that we could target and then fix the disease?
3:26 CA: But you're seeing now an opportunity to have a different kind of assault on these diseases.
3:31 What has changed?
3:33 SK: I think there are now three things that have come together just in the last one or two years that make it possible to understand such a complex problem like Alzheimer’s disease and other diseases like it -- and , at a high level, three areas, if you summarize it really quickly: measuring, changing and understanding.
3:53 And so measuring -- what that means for us is really single-cell sequencing.
3:58 So this is a technology that allows us to look at one cell at a time and take a snapshot of key dynamic processes in the cell which is the RNA expression of the cell.
4:08 So basically, RNA is like the language of the cell, and this takes a snapshot one cell at a time of what's going on inside it.
4:15 And then the second step, which is changing, we need to have the ability to change something very precise -- so changing one gene at a time stop it from making the RNA or changing it to upregulate the RNA.
4:30 This is the area that I’ve been working on now for 15 years -- CRISPR technology -- and as a field, we’ve made a lot of advancements.
4:38 And now we can do this across all the genes in the genome.
4:40 We can make these changes in a targeted way.
4:42 And it is really only possible very recently.
4:46 And then finally, of course, AI is at the forefront of everything, especially today.
4:53 But we’ve just seen over -- I would say really the last two years -- that ... it’s really working.
5:00 AI can help us understand these kinds of processes.
5:05 CA: So if I understand right, just as AI has cracked understanding human language, you see a possibility that AI can be used to understand the language of our own cells, RNA.
5:18 SK: Yeah, exactly, that's basically the core principle.
5:22 And for that, you need to be able to measure it and change it in this targeted way.
5:28 But as an analogy, the field was doubting this at the time.
5:31 I mean, even six years ago, people were not sure that you could really scale these large language models, just based on language and kind of predicting language, to actually build a conception of the world essentially, and at least approximate intelligence clearly pretty well.
5:52 So this is the key insight for the last six years, which is that a model can learn so much just from human language.
6:01 And similarly, we can apply that concept to RNA, which is basically the language of the cell -- especially the dynamic language of the cell -- because it's changing all the time.
6:11 It reflects what's happening to the cell, but also it reflects the cell's genetics.
6:15 CA: Is it approximately the same level of complexity as human language or much more so or less?
6:20 SK: It's hard to say, but I will say one key difference for me, and I think this is why AI can be so powerful for biology, is that human language is generated by humans, right?
6:30 So we understand it, right?
6:32 We came up with it.
6:34 RNA language, or the biological language, has evolved -- it was not generated by humans.
6:41 So it's basically impenetrable for us.
6:43 We can predict the left side: “to be or not to be” we know Shakespeare -- we can complete it.
6:48 On the right side , no human really could complete that, right?
6:52 But AI doesn't care.
6:54 CA: To try and crack it, I think you have to take the same stance of just getting huge amounts of data.
7:02 Talk about that process.
7:05 SK: Yeah, absolutely.
7:06 I mean, really what we learn, again, for large language models, is just they're very hungry.
7:11 They’re very data-hungry.
7:13 And really we’ve been generating data for these language models for thousands of years.
7:22 They're using all human languages that's been generated over generations and civilizations.
7:28 In biology, we don't have anything similar to that, right?
7:31 Especially when you're thinking about, we need these precise measurements kind of one cell at a time.
7:36 And we also need to know what actually happened to that cell because we're trying to build a predictive model, a dynamic model, that can predict how a cell will change when something happens to it.
7:45 And so we need to generate that data set.
7:47 And that's kind of really core to being able to build any useful model here.
7:52 CA: So give a sense of how you actually do this.
7:56 SK: Essentially, this is really combining those first two elements I was talking about, which is making a targeted change.
8:03 In this case, we're using CRISPR technology to turn a gene off or to turn it on.
8:08 And we're doing that one gene at a time for one cell at a time.
8:12 And then we’re measuring the outcome using single-cell RNA sequencing so we’re capturing what happened to the cell.
8:18 CA: So you do what you call a perturbation of the cell.
8:21 And then you measure the output.
8:23 How many experiments like that do you need to do, what's your plan?
8:27 SK: So our plan is to do at least a billion of these experiments.
8:31 So it's a lot of experiments.
8:32 (Laughs) Over the next four years.
8:35 CA: And you're not talking about, like in software, you're talking about a billion actual, biological -- SK: Yeah, they're all physical experiments.
8:42 I mean, I'm a biologist, an experimental biologist.
8:46 We're working with a lot of experiments in the lab.
8:48 And yet the way that we can do this is kind of using some tricks that make this much more scalable.
8:53 We’re not actually like, running a billion little individual reactions.
9:00 We're able to use kind of different bar-coding technologies to run these experiments in this bigger pools and then back out what happened what we did to the cells.
9:11 CA: OK, so if things work out as you hope -- I guess you’re already seeing evidence that it’s working out.
9:18 Once you gather that data, you're able to get from the model something truly amazing.
9:24 Talk about that.
9:25 SK: So just to give a sense of why I feel that we can do the billion experiments is we've done about 60 million experiments so far.
9:32 CA: You've done 60 million, right.
9:34 SK: So we feel pretty good we can keep going.
9:36 But yeah, the whole point of this is that we want to learn.
9:40 If I have this cell, and I make this change, what happens to the cell?
9:45 Really my motivation for generating this model is ultimately for human health.
9:52 And so for that, we can now have a disease state.
9:55 And importantly, this can be, for example, a certain cell in Alzheimer’s disease -- let’s say it’s an immune cell in the brain: microglia.
10:03 And we can measure, what that looks like, not just for one patient, but across many patients.
10:08 And this data is out there, so we don't even have to generate it.
10:11 And so we can see, OK, all these diseased cells, then we can have all the healthy cells, but again across people.
10:17 And then we can ask the model, OK, the model knows how to change cells, right?
10:22 So what intervention, what genetic change, what chemical change do I need to make to convert all the diseased cells across all the patients with the same disease back to the healthy cells?
10:33 CA: So that's an amazing sort of prediction.
10:38 Like, if you truly understand the language of DNA, the model can predict something that medicine has never known before, because the answer to doing that might be quite a complex series of interventions are needed for that cell.
10:50 It's not like you just give it an aspirin.
10:52 SK: It could be that it's a complex combination of things, or it's really just a question even of picking the correct one, right?
10:58 There’s, you know, 20,000 possibilities -- could be up or down to 40,000 possibilities.
11:03 And normally, the way this target identification and biomedicine works today is really this kind of guess and check approach.
11:11 So you have a hypothesis, one gene, then you're spending a few years on checking whether that's the right one, right?
11:18 So if you have 40,000 things to pick from, even if you just have to pick a single one, that takes forever, right?
11:23 And that's why we haven't cured these diseases yet.
11:28 CA: So what are you going to do with this model as you gradually refine it?
11:34 I mean, I understand, you think of this as basically a virtual cell, is what you’re creating -- it’s almost more than that.
11:41 It's like a universal virtual cell that researchers can, whatever cell they’re working on, use your model.
11:49 Talk about what you're planning to do with it.
11:52 SK: The whole point of it is that it is a universal virtual cell, which means that it needs to learn how to generalize to a new kind of cell or a new state of a cell, a new disease, for example, without having seen data, training data for that new cell type.
12:06 So that is a very challenging task.
12:08 And that's why we're really thinking hard about how to do these experiments.
12:12 But ultimately, I mean, this is what we're seeing here is, you know, the vision is that this is actually real.
12:18 So this is our state designer.
12:20 So we have already built our first model that came out eight months ago.
12:25 It's not very good.
12:27 So I mean to be clear, it is state-of-the-art, it's the best model at the time that was published.
12:32 (Laughter) But it has a really long way to go to be at the accuracy that I think it needs to be to be really useful.
12:43 But this is an interface that uses that model that we have today.
12:46 And so what you can do is you can say, “OK, I have this cell that I’m starting with.
12:50 And then, you know, I want to change this about the cell.” And then it spits out these different changes that you can make to the cell that are most likely to shift it the way you want.
13:02 CA: So you're not holding on to this yourself or licensing it to companies, you're making this generally available.
13:07 SK: That's right, yeah.
13:09 So we have a few ways that we really want people -- (Applause) People to be able to interact with it and also follow along.
13:18 So one is we're going to be releasing this tool later this year for people to try.
13:24 We'll give caveats like, this is not very accurate or this is going to be 20 percent accurate, but also we're going to iterate over the next four years.
13:33 We’re also hosting a “Virtual Cell Challenge” every year for the whole community.
13:38 We have 1,000 teams participating in the first one, and that's really to move the whole field forward, to get to where I think we need to be.
13:47 CA: So this is amazing.
13:48 The amazing work in your institute is really going to catalyze research worldwide because you're making this tool available.
13:55 Some people looking at that may go, “Well, wait a sec, isn’t that a little bit dangerous?” Some of the people playing with this may not have humanity’s best interests at heart.
14:03 What do you say to that?
14:05 SK: That’s definitely a question that I think, as you know, we got during the Audacious process.
14:10 And I think the key thing to keep in mind is that this is really just for human cells.
14:14 In theory, someone could build this kind of tool for a virus.
14:19 And I would say, don't do that, that's a bad idea, because yes, then you can absolutely use it to create something that would be dangerous.
14:27 But this really just allows you to shift human cells into a different state.
14:32 And I think that would be pretty difficult to abuse.
14:35 CA: And in principle, if a nasty virus did come along and this model is working properly, that's a way of giving us one of the quickest -- SK: Yeah, I mean, it will tell you, for example, how the virus is targeting this gene in this cell right now.
14:49 OK, we know what happens to the cell when that's getting targeted.
14:53 So yeah, it will absolutely help us to defend.
14:56 CA: So here's your team.
14:58 Tell us about them.
14:59 SK: Yeah, so Arc was only started in 2022 -- when we decided to launch.
15:06 2022 is really when we got up and running.
15:08 So it's only been four years, but we've grown a lot.
15:10 CA: So that's the before and after over four years?
15:12 SK: That’s just one year of growth -- the last year.
15:15 So we're over 300 people now.
15:17 And I think one thing we really wanted to be able to achieve with Arc is to bring people together from different disciplines and have AI and biology under one roof in one institute.
15:29 And we started this just around the right time when we could see what machine learning was going to mean for biology.
15:37 CA: Silvana, I got so excited to see the Audacious community get behind this and really help you expand this vision.
15:44 It's hard to imagine a bolder effort at really tackling what humanity needs and in making us all feel better about AI.
15:54 So someone here who's got in their family, they've got Alzheimer's or they've got heart disease or whatever, what would you say to them?
16:03 SK: I mean, I would say that I really think the medicine is going to transform for these kinds of diseases, right?
16:10 Maybe not in three months.
16:12 So you have to be a little patient.
16:14 But I think that within four years, five years, we will be able to have these models that are accurate enough to be useful.
16:22 And then it's a totally different way of doing biology.
16:25 It's not kind of, one hypothesis at a time, right?
16:28 A field like Alzheimer's can get really bogged down by just focusing on one dominant hypothesis that might be wrong.
16:35 And with these models, you can actually take a comprehensive data-driven look, you know, all the things that we could be targeting with a drug, what's going to happen with all of them, and then which of them is going to be the most effective one?
16:48 It's just a totally different way of tackling the problem that I think is so exciting.
16:53 CA: Silvana, thank you for your incredible vision, for sharing it with us.
16:56 Thank you, really, just fantastic.
16:59 SK: Thank you so much.
17:00 (Cheers and applause)