Podcasts

Ant Rowstron | What happens when you let an AI run a science lab

Listen on:

about the episode

There are AIs running labs with little human intervention right now. And they’re doing experiments that human experts would never try. Is this changing what and how science gets done? 

In this episode, we speak with Antony “Ant” Rowstron, who has worked with ARIA (the UK’s Advanced Research and Invention Agency) on their biggest bet to date. 

We cover:

  • How ARIA funded twelve teams to build AI scientists that can run an entire research process: generating hypotheses, designing experiments, and carrying them out without continuous human intervention. 
  • What AI scientists are actually achieving now, from personalized cancer vaccines to molecules that stimulate our own immune response to new viruses in 48 hours.
  • How we could train AIs on the tacit, hands-on knowledge only human scientists have.
  • How AI hallucinations might actually be useful for scientific discovery.
  • How labs and the role of human scientists will change as AI automates more and more parts of the research process.

About Xhope scenario

Xhope scenario

No items found.

Transcript

[00:00:00] Ant: You could be having a chat with an AI scientist who might well be able to design experiments. In the more advanced ones, it can then actually run the experiment. We've got a team at Liverpool. They've run their AI scientist, but at certain stages in their experiments, they've paused it, and they're going, "Okay, now, if we're as humans and we had these problems, what would we try next in order to make progress?" And they've also let the AI scientist models carry on. And it's interesting because at the moment the AI scientist is doing no worse, but not necessarily doing much better. But what's interesting is that it is doing different things.

[00:00:36] Beatrice: I'm very happy to be joined today by Antony Rowstron, who is the first ever CTO at ARIA, the UK's Advanced— what is it, Innovation and Research—

[00:00:46] Ant: Advanced Research and Invention Agency.

[00:00:48] Beatrice: Yes. Thank you, thank you. And you're the first CTO ever, but now you're Chief AI Scientist, as I hear it. And before that you were twenty-six years at Microsoft Research, but now you're responsible for this really big call at ARIA for AI scientists, which we're going to dive into today. So, for someone new to this, what is an AI scientist, and why do we want to create one?

[00:01:14] Ant: Yeah, so this is a really interesting question. When I first started looking at all the AI that was happening out there and all the terms that are used — AI for science, and so on — there's just so many terms being used. What really struck me was that there was an architecture emerging that was sort of three layers. You have at the top layer systems that can reason. So when I say "reason," they effectively try to create hypotheses, they ideate, they try to design experiments, and they're often based on large language models, frontier models. And then in the middle you tend to have these tiers, which are people who are working on models. So they are building a tool — AlphaFold would be an example. And so they tend to be building models that are predictors of things. So you have a dataset, you model it, and then you can predict something about whatever the dataset was trying to model. Extremely important. And in some senses you can think of that top layer as needing to call and use that second layer. And then I think underpinning it all was the emergence of more automated labs and the idea that doing everything in silico was really hard — doing everything on a computer, extremely hard. And wouldn't it be useful to have automated labs that could both be in the loop and, at the end, doing verification? So you can create data when you need it, you can check an experiment, you can check a hypothesis. So the AI scientist layer, to me, is that top layer. AI for science, to me, is very much that middle layer, and then it's underpinned by automated labs and the datasets we require to make them work.

[00:02:50] Beatrice: So, so then your call was focusing on that top layer, like the hypothesis layer.

[00:02:54] Ant: So we were asking for people who actually had either that full stack — but they had to have the top layer — and ideally they also had to have labs that they could demonstrate that top layer was interacting with. Many of them are using the middle layer as well, so they have tools which are similar to the AlphaFold-like tools, and they're incorporating them. So the top layer can call out to the tool when it wants to, the top layer can call out to the lab when it wants to. And so that was the thing we wanted to understand — where we were on the boundary of that. So it would be unrealistic to say that exists completely today. So the question is, where is the boundary? What can we do? What can't we do? Where do we need to push? What's already happening?

[00:03:39] Beatrice: And if we dive deeper — how does it actually work? What does an AI scientist actually do? Do you have a concrete example of one?

[00:03:48] Ant: Yeah. So, for example, you could be having a chat with an AI scientist. Think about it like — if you were interested in exploring a space, you could have a chat with it. It could generate ideas that you might want to explore. If you said, "Yeah, I like that idea, let's explore that more," it might well be able to design experiments. And in the more advanced ones, it can then actually run the experiment. It can look at the results, evaluate the results, and decide whether the hypothesis was wrong or the experiment was wrong, and re-loop around that. So I've seen systems where — and sometimes these are more specialized into particular domains, sometimes they try to be a bit broader. In effect, some of the teams we fund are fairly advanced. One of the most impressive ones was when I got to a meeting and they said, "We've got a system that can effectively generate quantum dots on demand," which is actually quite hard. So this would be like a tool where you could say, "I'm really interested in getting a quantum dot, and I want it to have these properties." And we started the meeting, and they asked us a number of properties that we might want a quantum dot to have. Quantum dots are a bit like an LED — they emit light under some circumstances. So you can say what wavelength of light you want it to emit, and things like this. And by the end of the meeting we went into a lab, and they had a tray of quantum dots — each tray corresponding to a different person at the meeting, and what they'd asked for as the characteristics of it.

[00:05:24] Beatrice: And, yeah, maybe that's interesting — could you talk about one or two of the projects that you ended up funding?

[00:05:31] Ant: Yeah, yeah, I think there are some really fun ones. One of the ones which — and they span the whole spectrum. We funded about 12 teams around the globe. So we wanted to really see what the full frontier was like. About half of them are in the UK, half of them are outside the UK. One of the UK ones I'm super excited about at the moment is based at Oxford University, and it's looking at personalized cancer vaccines. You know, the basic premise is: when you have a cancer, if the tumor's removed, they need to be able to see whether there are still cancer cells, and what type of cancer cells. And they'd like to be able to take a blood sample and measure what type of cancer you've got — and there are about ten thousand different cancers you can have. And then they want to create a personalized vaccine which protects you, or tries to kick your immune system into attacking the top twenty they've identified from your blood sample. So you get this personalized vaccine, which then attempts to stimulate your body's own defenses to attack the cancer cells that are remaining. And that's really interesting because it's actually enabled by all those layers. They're really using AI-for-science models, they're using an AI scientist above it, and they've got automated labs which are allowing them to accelerate the progress they're making. So I think that's one of the really exciting ones. I think another really exciting one was a team we went to visit in London, AminoAnalytica — a very small outfit, three people, basically. When they started, they'd been looking at COVID and creating antigens that could bind to COVID, which would be used in lateral flow tests and things like that. And they wanted to automate the design of those antigens, and automate the testing of them, and find suitable ones. And we went there just as the hantavirus was taking off. And they sat there and said, "We could probably create an antigen for hantavirus in about forty-eight hours using the tool chains we've got." And I was like: wow, that's just really crazy. And their view is, if they can accelerate everything, they'd be able to come up with a lateral flow test in three weeks for a new disease, a new virus, or whatever. And I think that's another one that was accelerated by having an AI scientist layer, and also AI-for-science models that they were using, combined with labs to then test and refine the data, allowing them to loop around really quickly.

[00:08:15] Beatrice: Much faster feedback loop.

[00:08:16] Ant: Yeah, yeah.

[00:08:17] Beatrice: And, yeah, you said you funded twelve teams, but there were two hundred and forty-five applications, which is the largest response ever.

[00:08:24] Ant: Yeah. At that time it was the largest response ever. I was super surprised. And we'd also set this up to run as the fastest thing ever — one of the things about these tools is that everything's about velocity, so I also wanted to do it quickly. So we wanted people to propose, we wanted to sign them quickly, we wanted to get them going. I can't remember the exact date we opened the call, but it was around September. And the idea was to have them all signed by Christmas. So of course, with two hundred and forty-five applications, it's almost a "success disaster," because we had something like three weeks to select which ones of those we funded. And we wanted to read them all, we wanted to look through them all. And then we managed to select twelve that we were really keen to fund. And most of them were going by January, great — many of them started on the first of January, or the second, whatever. Yeah, so that was amazing.

[00:09:14] Beatrice: That's a good job, especially within a government agency.

[00:09:17] Ant: Yeah, well, that's ARIA — yeah, we're very fast when we need to be. It's good.

[00:09:20] Beatrice: And, as I understand it, you also structured it as nine-month sprints. Why — how come you chose that?

[00:09:27] Ant: Well, actually, we only gave them nine months in total. So, one of the things about this was: we didn't want to fund people who were creating systems, we wanted them to already have a system, and we wanted to know what it could do. So we decided that rather than doing the usual two-year or three-year term, we'd just do it as nine months. If you have these tools, they're supposed to allow you to go faster — so let's try to get it all done in a short window. And for each of the teams we asked them to do two things: something that they thought they could do, and something which they thought would be a real challenge — probably wouldn't be able to do, but they'd like to have a go at. It was also unusual because ARIA has "opportunity spaces" — areas where we think there's ripe for disruption — and we only usually run programs inside those opportunity spaces. But for this AI Scientist call, we said you could actually select what you wanted to solve, and it didn't have to be in one of our opportunity spaces. Now, most of them — our opportunity spaces are quite broad, so most have ended up doing things inside our opportunity spaces, but one or two are doing things outside. And that's fine, because we really wanted to see the how — how they were doing it, and how well it worked — rather than necessarily the what. The what was kind of the bonus, which is unusual for funding. Normally funding is about the what you achieve, not the how you did it. Whereas we wanted to invert that — we were super interested in the how, rather than the what.

[00:10:55] Beatrice: And is that because you think that will inform your work so much? Or —

[00:10:59] Ant: Why — yeah, I think it will inform ARIA. It will help us understand how to think about how these things will operate in the future. And I think one of the things you notice — I noticed it as a computer scientist — is that it's often hard to adopt new tools. It's often hard to understand how to incorporate new tools into a workflow to accelerate what's going on. So it's also helping us think: okay, the successful teams that are really maxing out what they can do — how are they using these tools? And can we then use that to help other creators in our ecosystem think about how they might access these things to accelerate what they're doing?

[00:11:36] Beatrice: And do you have a sense of — what do you think is likely that they will achieve in nine months, for example? Like, if you take the cancer vaccine.

[00:11:46] Ant: Yeah, I think — so, I think they'll get to the first stages. They will get evidence. Most of these things are going to get evidence of something, rather than complete something in its totality, as it were. So each of these will demonstrate something that would be, you know, a step forward if they can do it. But it's not the end. So for the cancer vaccine one — and it's super interesting, actually — they've now got a mentality of going fast. So, in a sense, they're talking about how they get from where they'll be at the end of the nine months to actually being in humans. So they'll have evidence that these vaccines can be created, evidence that they may work, but they won't have human trials or anything like that. But the interesting question is how quickly can they get from the point when they think they've got something to the point when they can try that. And it's fascinating, because I think there's a mindset shift in velocity — the rate at which they think they could do it is super interesting. And I see this across all of the programs — all of our creators in the program are really thinking, "Okay, how do we accelerate to go faster? How do we exploit that speed and keep that speed going?" Which is interesting.

[00:13:06] Beatrice: Yeah, that's kind of the — when I was reading up on this for this interview, the term "more shots on goal" was what came up with the AI scientist thing.

[00:13:17] Ant: Yeah, well, I don't know whether they're shots on goal or just — it's an interesting question. At a high level, you go faster. If you go faster, it means you could perhaps have more shots, and hopefully some are more likely to be on goal. I think also — it's interesting — we've got a team at Liverpool who are really, really interesting. What they've done is they've run their AI scientist, but at certain stages in their experiments they've paused it. And they've taken a snapshot of the universe at that point in time, and they're going, "Okay, now, if we're as humans, and we had these problems, what would we try next in order to make progress?" And they've also let the AI scientist models carry on. And it's interesting because — the results are still coming in, we're only four or five months in — but at the moment I'd describe it as neck and neck. We're not — the AI scientist is doing no worse, but not necessarily doing much better. But what's interesting is it's doing different things. So talking to the professor who runs it, he's like, "Well, these are the things I did because thirty or forty years of experience tells me that's what we should do — but it did something that I wouldn't have done." "And actually, it's done as well as what I tried to do." And in fact, interestingly, sometimes the things that are the obvious things to do with all that experience are the things which, when you look at the result, don't provide any benefit at all — and have been completely avoided by the AI scientist. So I think of it as an interesting one — it's not yet beating— but it's doing it differently, which I think is super valuable, because probably that means you're going to see things you wouldn't expect. We have another one looking at endometriosis, and they're looking for targets that you might attack — or use to, and it's interesting because they found targets that experts were like, "We're not sure we would have gone for that." So they're pursuing them and trying to see: are they actually targets, and are they good targets to target? So this is interesting — it starts to surface things you wouldn't normally try. Which I think is super exciting.

[00:15:43] Beatrice: It's super interesting. It's like that canonical AlphaGo example, where the AI playing AlphaGo made a completely different move than a human could have ever thought of, and so on. 

[00:15:55] Ant: Yeah. And I think we're starting to see bits of that come through, which — I think we're a long way from just setting a system off and saying, "Go find me a solution to X." But, to me, that's partly the power of all this — the power of being able to explore things you wouldn't normally think about. That's partly the power of these tools. And I think, when you think about how they're operating, they're very much "human on the loop, human in the loop." And I think they work in partnership with human scientists in a way that will probably yield things we wouldn't expect.

[00:16:33] Beatrice: Well, I think that's actually an interesting question too: how do you expect the role of the human scientist to change in the next ten years or so?

[00:16:41] Ant: Yeah, no, I'm often asked this question. Now, the really interesting thing is I'm not a scientist — well, I'm a computer scientist. Many people argue whether computer science is computer science or whether it's something else, but it's super interesting, my experience. I'm probably one of the luckiest people, because these coding tools are progressing at such a rate that they're unlocking things. I spend a lot of my time creating and trying things out that I could never do normally. For the last ten or fifteen years of my life, I spent most of my time managing, looking after people, providing feedback. Now I can actually have ideas and try them out really rapidly. It's really empowering. And I jokingly think I'm probably the most innovative I've ever been in my career. So in that sense I think scientists will get to the same spot, where instead of worrying about getting up at three a.m. to go to the lab if you're a PhD student, instead of thinking about all the nuts and bolts that have to get done to explore an idea, these tools will enable you to just rapidly explore spaces. So in a sense I feel they're going to be tools that just amplify what we can do. As a computer scientist, I find it — it's difficult to explain to a non-computer-scientist. But the funny thing is, I meet lots of people at my age and stage, people I worked with twenty or thirty years ago, and we all talk about the same thing. We're sort of re-empowered. Because you don't need to sit there for three or four weeks to write a huge amount of code. You can do it in a couple of days interacting with Claude Code or one of the other tools. Very, very empowering, to explore space and ideas.

[00:18:42] Beatrice: Yeah, that actually makes me want to ask about some of the challenges around this. Because one critique I've heard in relation to AI scientists is that maybe we're extrapolating from computer science, and that it's easier for an AI scientist to succeed in in-silico things. But in a wet lab, for example, when you're dealing with physical reality — are we forecasting from the easy cases to the hard ones, or do you think we will actually get there?

[00:19:22] Ant: So I think this is a really good question, and — in some senses this is almost the fun of the field, because where will we get, and what will be possible? It's super interesting. Coding is a very — there's a lot of examples, and there's also a very easy way to check whether things are effectively running: are they correct, are they compiling. You get a lot of feedback dynamically, which these tools use to work out whether or not they're on the right track. Science is obviously very different. And it's interesting, because if you talk to a biologist, it turns out a lot of discoveries in biology were serendipitous — they weren't the hypothesis being tested, they were a byproduct, and knowledge and understanding of something they were seeing allowed them to go off on a new angle. So it's super interesting whether these tools will be able to do what's needed to really create new ideas, new things. What I would say is — it's an interesting question whether you can actually create these tools to be more exploratory. At the moment, a lot of people are trying to create tools that take all the knowledge, search it, find things, and push forward. There's an interesting inverse question: if you get a really good understanding of what knowledge exists, you just look for the holes, and then start to randomly explore the holes, rather than very carefully steering towards one particular thing you're trying to get to. So, in summary, that's a rather long-winded way of saying: I think there's a lot that we don't know if we can do. But there are also a lot of early signs that we might be able to do them, and that's why I think it's interesting to keep thinking about it.

[00:21:20] Beatrice: I mean, it also brings us to the question of the wet lab, or the lab, actually. What do you think — we spoke about how the human scientist's role might change. What will the lab look like in five to ten years?

[00:21:38] Ant: So I think this is a really interesting question. I remember when cloud data centers first started appearing, and I remember going to a very early cloud data center and being completely awestruck at how different it was. And at the time, a lot of people on my team were looking at building things that would go in what we'd call an enterprise data center — a data center a company would have at the top of their building, or in the basement. And when I went round a cloud data center, I left that day realizing the world had changed — that almost everything we were doing wasn't going to be relevant in that new cloud era: the hardware we were building, the way we were thinking about systems. When you thought about them in this new world, they just didn't make sense. At that stage, we pivoted the team to look at technologies that could really help grow the cloud. When I look at labs, I feel they're a bit like the era when you had a personal laptop that you owned and liked to use. And I think these automated labs will become much more like servers — the equipment will become designed to support things at scale, at speed, and to be shared and used. Effectively, the most cost-effective way of using a piece of equipment is to use it nonstop. The closer you get utilization of anything to a hundred percent, the cheaper it is to run. So I think these labs are going to evolve hugely. At the moment they're set up for a way of working that's quite ad hoc — some pieces are automated. I go and visit labs, and it's kind of crazy — people buy this really expensive automated instrument, and then remove the front panel because they can't get something into it automatically unless they take the front panel off. They can't get information out of it. There are examples where the temperature is shown on the screen, but you can't programmatically access it. Things like this. So there's going to be a change in how we think about hardware, similar to the change from data centers to cloud data centers. So I think labs are going to change a lot. The other interesting one — I was out visiting Edison FutureHouse, one of the teams also involved in this — they were saying that part of the challenge is, you've got a lot of labs which are like the Ford of labs, and then you've got an F1 outfit that wants to run really specialist experiments. And this is the interesting point: science is very broad, even within a domain like wet labs that support biology. There's a lot of diversity in what you can do in those labs. Thinking about which ones to start with, how to expand the scope for automation further out, I think is also super interesting. And I think we might even see new instruments that don't yet exist. I was at a talk the other day by Tom from Adamo, a design company. And he was highlighting to us that a new instrument has almost underpinned every great breakthrough — the instrument was created before the breakthrough happened. And I think, as we evolve these labs and stop thinking about whether humans need to necessarily be able to interpret all the results that come out — can you get away with having other layers that interpret results and then present them to humans? MRI scanners are interesting: the data an MRI scanner produces is so complex that a human can't just interpret it. You have to have layers that subsample it, to interpret it, to then display it in a way a human can interpret. I wonder if there will be more instruments like that in the future.

[00:25:34] Beatrice: I didn't know that about MRI scanners, but that's amazing.

[00:25:30] Ant: Yeah, yeah, it's kind of interesting. MRI scanners are kind of crazy.

[00:25:34] Beatrice: Oh, they're good though, I'm very happy we have them. But it brings me to another question, which is one of the potential challenges — something that often comes up when you read about how science is done — is that experimentally there's often tacit knowledge: knowledge that you learn by being in the lab, doing the experiment, watching someone else, and that's not easy to convey in language. And that might be a challenge for getting an AI in on that layer. How do we get an AI scientist to be good at that, or how do we get around that challenge?

[00:26:16] Ant: So that's another really interesting challenge, and, again, there's no quick answer. I think one of the most interesting things I've seen at the moment is that there are a few outfits around the globe — some startups, some FROs and things like this — that are starting to do things like record what people do in lab spaces. So they're recording from the human's perspective, or from above the bench, or both, and using that to try to understand what the human is doing, to observe things that might become that tacit knowledge — if you shake the test tube at the right point, it works well. Well, they capture the fact that the test tube gets a little shake. So I think that's a super interesting approach — can we try to capture some of that information? I think also, there will be a world where, if we build these systems correctly — yes, they'll make mistakes at the start, but as they understand more, they'll in effect begin to codify the tacit knowledge, because they'll realize something didn't work and they need to do the experiment differently to get it to work. And there's a world where that could even be better for us, because I think one of the problems at the moment, in many scientists — not all, but some — is that reproducibility is really hard. And, as a computer scientist, that's kind of crazy — if I have a piece of C code, I expect it to run on every processor the same way, I expect everyone to be able to use it. I'd find it weird if it just didn't work. So I think, in some senses, capturing some of this tacit knowledge and being able to encode it in the experiments directly will actually help reproducibility a lot, because in a sense it won't be about knowing that you shake it — it'll just say in the instructions: shake it. And so, again, great question — we'll have to see how it pans out, but there are early ideas that seem to be pointing towards what we might at least be able to capture of that knowledge over time.

[00:28:26] Beatrice: I mean, just filming someone in detail, breaking it down, seems like it should help.

[00:28:31] Ant: Yeah, and it's an interesting one. So I went to one of the startups in Boston, and for the first time in my life they got me to pipette, and they filmed me doing it. And I was there with a colleague of mine who's actually a biologist, which is really unfair because he's worked at the bench for many years, before working in our area. And I actually thought, when I got to the end of it, I'd done quite well and followed the instructions really carefully. Really carefully. And I had someone beside me going, "No, wrong pipette," and "Wrong scale." I thought I'd done really well at this. And then we got the scores back, and out of two or three hundred people who'd tried it, I was worst. And to make it even worse, the guy I was with was actually the best. So it just shows that it's interesting — even when you think you're following the recipe exactly — an assay, sorry, or a protocol exactly — you've done it badly. So, yeah, that was a really interesting one. So anyway, they should definitely learn off my colleague, not off me, when they give it to me.

[00:29:32] Beatrice: Well, I'm sure — it seems like as long as we can keep training them, that's going to work. Another question: do you worry that AI scientists, if they're all trained on the same data, that we risk losing some of the creativity that's needed, or some of the outlier bets you need for breakthroughs, because we homogenize the direction of science? Is that something you've thought about?

[00:30:07] Ant: I worry a bit about that. I mean, the other part of that is thinking out of distribution — can you think out of distribution? And if you can't, do you just homogenize on the distribution, as it were? I think it's a really interesting question, and I think we need to study it more. I think this comes back to being intentional about exploring — in some fields, exploring more in a random way. Not random, but rather than saying "this is exactly what I'm going after," try to think about how we create tools that are able to explore spaces. And I think that might be one of the things — you almost get serendipitous discoveries by covering spaces you're interested in, rather than trying to pick a point. You use the velocity these things can go at to cover a set of points in a way that would be unfeasible at the moment because of human bandwidth and time. The number of PhD students who can come in at three a.m. is limited, whereas if these things are more automated, you're able to cover a lot more space a lot more quickly.

[00:31:17] Beatrice: So, another potential limitation I was reading people think of is that a lot of — the AI scientists currently are based on LLMs, right?

[00:31:29] Ant: Yeah.

[00:31:29] Beatrice: And so these systems can be a bit unreliable, sometimes even make things up, because they're trained to predict the next token, sort of. And so for the parts of science where you need a really precise, repeatable answer, you might rather use some sort of traditional software to get it right every time. But then you have the creative parts we also talked about, like coming up with a new hypothesis. Maybe you also can't trust an AI LLM there, because there's no good way to check if it's right. So, how do you think about that? What's left in terms of the value of the AI scientist?

[00:32:14] Ant: This is another really good question. So, let's take the hallucinations topic first. I think partly one way of thinking about this is: when we first got LLMs, you were treating them more like an oracle. You'd assume they had all the knowledge inside them, and you'd say to them, "Tell me about X," and it would get it wrong, and you'd go, "Okay, that's a hallucination." And it didn't know it was wrong, and it said it very confidently. And people felt this was not a great experience. I think one of the things that's changed — I might put it like: that's like expecting to go along to one person and having them remember an entire library, and just turning up to that one person and going, "Can you remind me what Shakespeare said in Act Three?" and that one person just sits there and tells you the answer. If you think of how a lot of people are starting to use these tools now — and in AI science this is really true — what they're saying is: actually, this person is going to be really good at finding the right book in the right location in the library. So you walk up to that person and say, "Like a librarian, I desperately want to know what William Shakespeare said in Macbeth, in Act Three," and they're able to walk almost directly to the exact spot in the library, pull out the exact book, and show you the exact text. Now, in science, what people are doing is starting to create these knowledge graphs — it's called grounding. You're not asking the LLM just to recall things. What you're doing is asking it to use some of the knowledge it has to think about how it would find more information from knowledge graphs and other forms of data storage that would allow it to think and reason about an area. It turns out that when you ground these systems, they tend to hallucinate a lot less. So it's a bit like — if you're human, and I asked you a question like, "What's the population of London?" You could probably give me a number, it would be a bit of a guess, but an educated one. If I said, "You're allowed to look in a reference book — what's the population of London?" — you'd know probably which book to pull out and which number to look at, and then you'd report back that number. So it turns out that when you ground these things, you give them context which helps them — they have knowledge, but the way they're applying it is different from just being asked to regurgitate things. So on the hallucination side, if you look at many of these modern systems, because they're a lot more grounded, they hallucinate less. Doesn't mean to say they don't hallucinate.

[00:35:04] Beatrice: That makes a lot of sense.

[00:35:05] Ant: Yeah, they’re sort of grounded when they do. Now, if we go back to the biology point, and the fact that many discoveries are more by chance than as a byproduct of trying to do something else, it may be that in some spaces hallucinations are not actually that bad, because that's a little bit like somebody doing something and going, "Oh, wow, that's not what I was expecting, but actually that's really interesting." Perhaps hallucinations are a way of getting some of that noise you need into the system, because it just thinks, "Oh, this should be true — oh, it's not," or "Oh, it is, but that's not what we were trying to answer at all — how interesting." If that makes any sense. And what was the second part of the question?

[00:35:53] Beatrice: It was the hypothesis part.

[00:35:55] Ant: Yeah, out-of-distribution thinking. Yeah, that's really interesting as well. So, at the agency, one of the things we were thinking about and discussing a lot was: how might we actually start to measure this? Because this is one of those questions everyone always asks. So we've got this really interesting — what I think of as a small suite of benchmarks. What the benchmarks do is simple: it creates a small virtual world with particular properties completely different from Earth. We have a little world for ecology, a little world for biology, a little world for chemistry. You can think of it almost like a box experiment you can run in this different world. You have instruments that act as actuators within the world, and instruments that act as sensors to read data out. So any knowledge you have about how, say, gravity works on Earth is no good, because in this other world it doesn't exist. And then we created a very simple AI scientist, and we basically say to it, "You've got to work out the laws that govern this world." And we don't even tell it what subject it is. We don't say this is a biology world, or a chemistry world, or an ecology world — we just say, "Here are some instruments," and give them names. It's true that probably the names hint at the space they could be — so it's like "mix" and "heat" and "stir," things like that, as instruments they can access. And then we see whether the LLM can actually work out the rules, and we do it by having it generate hypotheses, design experiments, run them with the instruments, look at the results, and keep looping round, learning what it's learned, until the point it decides it's learned the laws and stops. And what's interesting is — some of the more modern models, some of the more recent ones, GPT-5.5, the new Gemini 3.5 Flash — those models are actually able to extract quite a lot of the rules from these worlds. Which is sort of an interesting — because they won't have been trained on these worlds. These worlds have extremely bizarre rules designed to be completely atypical. So does that suggest that they can, at least at some level, create hypotheses and generate experiments to prove those hypotheses? It's an interesting starting point. Now, we'll probably release something on the web in the next day or two, but we're hoping it'll seed a bit of discussion about what other things you could do. So it's not testing recall — it's not testing the ability to recall some particular fact, like what's gravity. We're not asking it that. And in fact, recalling what gravity is won't help it do the benchmark. It's more about: can you design an experiment given some tools and instruments, and can you iterate? And we've learned a few things along the way. It's interesting — if you ask it to come up with a single hypothesis, it doesn't do super well. If you ask it to create a set of hypotheses, and then have something else that selects the hypothesis given what the world looks like and where you're trying to get to, it tends to do better. So that's the interesting thing already — just asking a single line like "Come up with the next hypothesis, come up with the next hypothesis" doesn't seem to do as well as saying, "Come up with a set of hypotheses, and we'll work out which one we want to try to move forward with."

[00:39:27] Beatrice: Is this published, or is it—

[00:39:29] Ant: We will be — we'll be releasing hopefully in the next few days. It's super — I think it's super interesting, and we were going to call it a benchmark, and then we didn't call it a benchmark, because it's more of an exploration. Because I don't think it's actually a conclusion — it's more of a "this is an interesting path, and where can it go" kind of thing.

[00:39:47] Beatrice: Well, we'll make sure to link to it, then, when we publish this episode.

[00:39:52] Ant: Fantastic, thank you.

[00:39:56] Beatrice: Someone young is listening to this podcast, maybe, and they feel really excited and want to engage. Do you have any career advice for them?

[00:40:05] Ant: Wow, yeah. I think — one of the interesting things is, I think we are at a moment in time where things could change a lot, or they might not — one never knows. If you're really interested in this space, I'd suggest thinking a little about this: there are already people publishing bits of kit you can quite easily build to run experiments at home. You can link them up to your own setups, to try to do experiments on your desktop, so to speak. I'd suggest looking online for some simple — if you search for "AI and science" and "kits" and things like that, you'll find people creating quite interesting little setups you can experiment with. The other thing I'd say — and I'd say this to anyone young today — is that it's super important to be broad in what you think about. One of the things I've always done in my career is, every decade, try to change track slightly, try to do something different from the previous decade. Sometimes I manage to escape it more or less, depending on how I've done it. I think the ability to apply oneself in a different area and adapt to that change is going to be a really key skill in the future. So, in a sense, we often think of breadth versus depth. It's an interesting one — I think the future may favor those who can be broad more than it favors those who are extremely deep in a single area. And I think — it's interesting. I talked about how I was using some of these coding tools — it's super interesting. Over Christmas I set myself a challenge to write a paper. Now, I haven't written a paper on my own in thirty years, I've always had collaborators. And I also chose a topic where I'm not a domain expert, but I could understand it well enough to know whether it was going wrong. So it's a bit like: perhaps I can't write or read a language, but I can speak it. If I can then translate what's been written into the language I can speak, then I can know whether it's going correctly. And I was super impressed — in two weeks I actually created a system, started running it at some scale. It worked well. I wrote the paper. And it was in a space where normally I would have had collaborators with me who would have just been providing guidance — they wouldn't have been doing it, but they would have said, "That bit's not right," or "That doesn't look right from my perspective as a security expert in computer science," or whatever. Whereas, in a sense, I was able to use my own knowledge in those spaces, plus using coding tools and LLMs and other such tools, to help build a scaffolding in the space where I wasn't the domain expert. And I think that's what life's going to be like — there's going to be a lot more leveraging these to expand out, rather than just worrying about being extremely deep in a particular area. Does that make sense?

[00:43:06] Beatrice: Yeah, just like having a lot of great tools at hand so that you can just be a manager of many things.

[00:43:13] Ant: That's right. And I think people in the future — there's a significant chance that a lot of roles in science and things like that will be more about being able to use all those tools, and being less siloed. Science is so siloed. If you go along to a wet lab, it does a particular type of biology, and then you go along to another lab and it does another thing, and they don't — it's really weird, they operate independently. And yet you can see a world where the interaction between those areas might be so valuable. Or, you take a materials lab and put it beside a wet lab — what happens when you do that? Well, we don't know, because at the moment, in most places, you have a building with materials labs and a building on the other side of the city with wet labs, and there's no real interaction. And yet, fundamentally, there might be really exciting things that could emerge when those things come together.

[00:44:06] Beatrice: Yeah, yeah. I heard a quote from Tom Kalil — he was probably quoting someone else — which was, "The world has problems, but the universities have departments."

[00:44:16] Ant: Yeah, no, that's a great quote. I love it, exactly. And even within a department — the number of universities you visit where you have, you know, "X, Y, or Z lab," you have "Ant's lab"— and they do "ant stuff." And it feels a bit like, in this world, trying to break those barriers down, and thinking, okay, can we use these tools to help accelerate — even coordination between those—

[00:44:43] Beatrice: Translate between, yeah.

[00:44:43] Ant: Yeah. I used to lead a very interdisciplinary, multidisciplinary team at Microsoft, and I think for the first two or three years, when we started hiring physicists, mechanical engineers, all sorts of different skill sets, it was really complex, because everyone spoke a different language. And I think trying to even get to the point of speaking the same language — well, the great thing is that a lot of these tools can help translate. And I think that's why they'll enable so much in that sense.

[00:45:10] Beatrice: Yeah, I mean, that's definitely the case between different sciences, and just different communities at large — this translation piece.

[00:45:20] Ant: Yeah, yeah.

[00:45:21] Beatrice: Yeah. Do you have an existential hope vision for the future? If you think about the future, what are some of the things you'd really love to see in it?

[00:45:30] Ant: Well, I'd love to live in a world where, effectively, people didn't die from diseases that should hopefully be treatable in the future. I've already talked about the Oxford cancer work — I just think there are so many diseases and illnesses that people die from. And fundamentally, part of my hope about this whole AI-for-science space is that, in twenty years' time, no one dies from some of these things. And I think the velocity — the speed increase these tools can bring — could really help us.

[00:46:10] Beatrice: And if you weren't doing this, what would you be doing?

[00:46:13] Ant: If I wasn't at ARIA, what would I be doing? Sat on a beach, perhaps — where we're heading, I don't know. I think the other thing is, I probably always secretly wanted to do a startup, and I've never really had the guts, or perhaps the idea, or the team — I don't know what — to do it. But perhaps, in a parallel universe, I wouldn't have gone to ARIA. Instead I would have left Microsoft and done a startup.

[00:46:38] Beatrice: And the last question — what's the best piece of advice you've ever received? Or just a really good piece of advice you've received?

[00:46:49] Ant: So, I think one of the things that had the biggest impact on what I've actually chosen to think about technically was, really early on, some guidance about thinking about how the world will be in, say, ten years, and making sure that what you're working on today fits into that world. A lot of people who work in research tend to focus on the here and now, kind of ironically. So you build great tech that works really well for the world you're in now, but in a decade's time that will be different. I learned that lesson the hard way. Back around 2000, we were working on technologies to do with peer-to-peer systems and how systems could communicate with each other. It was really exciting, I was innovating at Microsoft, and we were looking at everything Microsoft had been doing in this space. And a very senior person made the observation that, if you solve for the problem of bandwidth — which is what we were doing, networking bandwidth at the time — in a decade's time bandwidth's going to be plentiful, and we'll have all this legacy complexity, and everything else we were proposing, which was fairly complex, would have no value. We'd be better off just thinking about what we're going to do in a world with lots of network bandwidth, and not worrying today about how to handle the fact that network bandwidth was relatively expensive. That was a really hard lesson, because I'd spent several years working on something, as had everyone else around me — there were about thirty or forty of us in the room who'd all been working on these kinds of technologies. It's a very hard lesson, but one that has stuck with me ever since: always, I think, if I fast-forward a decade, would this problem I'm trying to solve still really exist? And if it won't — if the root cause is just fundamentally going to change — why solve it? Focus on things that are going to make a difference for the longer term. So I think that's the piece of advice that's impacted me the most.

[00:48:51] Beatrice: Yeah, that's a great piece of advice, I think, to end on. Thank you so much, Ant, for coming and talking about this, and for doing this work. Thank you.

[00:48:59] Ant: Thank you very much. It's been great, thank you.

Read

RECOMMENDED READING

Organizations and projects

  • ARIA (Advanced Research and Invention Agency): The UK government's high-risk, high-reward research funder, which is running he AI Scientist program.
  • ARIA’s AI Scientists project: ARIA's page on the AI scientists call and the twelve teams funded through it.
  • The Cancer AI Scientist Project: The ARIA-funded project behind the personalized cancer vaccine work, aiming to unify tumor-immune AI models, lab automation, and sovereign compute into one platform for discovering cancer vaccine candidates.
  • Amina: The ARIA-funded project behind building an AI scientist meant to compress pathogen diagnostic development from months to days.
  • Edison Scientific (FutureHouse): The AI-scientist company Ant references when discussing how automated labs and instrument design are evolving.

Tools and technologies mentioned

  • Albert: An AI Scientist Benchmark : Ant's benchmark of five synthetic worlds with invented physical laws, built to test whether AI scientists can genuinely discover new rules through experimentation rather than recall known ones.
  • Stress-testing AI Scientists in parallel universes: ARIA's writeup on the Albert benchmark, with more detail on what it found across different frontier models.
  • AlphaFold: Google DeepMind's protein-structure prediction tool, cited by Ant as the archetypal example of an 'AI for science' model that an AI scientist can call on.
  • AlphaGo and "Move 37": DeepMind's Go-playing AI, referenced in the conversation as the canonical example of an AI system finding a move no human would have considered.
  • Claude Code: The coding tool Ant mentions using to write software in days rather than weeks, part of the broader productivity shift he describes.

To learn more about key concepts mentioned in the conversation