Podcasts

Anthony Aguirre | Tools or Agents? Choosing Our AI Future

Listen on:

about the episode

As AI systems become more capable and autonomous, staying in control of what they do is increasingly difficult. But the AI we need to solve problems like cancer doesn't have to be one we can no longer steer.

In this episode, we talk with Anthony Aguirre, cosmologist and co-founder of the Future of Life Institute, about how we can design AI to amplify human capability rather than substitute for it.

We cover:

  • The economic driver Anthony sees behind AGI: largely not scientific breakthroughs, but capturing a share of the global labor market
  • What "Tool AI" means to Anthony as an alternative to AGI, and why he thinks it can deliver most of what we want without replacing people
  • Whether Tool AI is stable: the tension between staying in control of AI and the ease of completing tasks
  • How legal liability for AI agents could quietly steer the industry toward more controllable systems
  • Two concrete ideas for transformative AI tools we could build today to improve democracy and the information landscape

This episode is part of our AI Pathways series, where we explore the choices we can still make about how AI gets built.

About Xhope scenario

Xhope scenario

No items found.

Transcript

[00:00] Anthony: The bad news is that part of the incentive structure for AGI is not curing cancer, producing immortality, or even producing the super beings that are going to solve all of our problems. It's replacing people. The real economic driving force behind AGI is that it is the thing that allows you to not have to re-scramble all the tasks in your company, not have to figure out how to automate certain things and have the right controls over them, but to just say, "No, I'm not gonna hire this person, I'm gonna hire this AI system instead."

As you imagine very powerful AI technologies, there's a level at which we are gonna have to make a decision. Do we want them to be our tools or do we want us to be their tools?

[00:39] Beatrice: Thank you so much for joining, Anthony.

[00:41] Anthony: Pleasure to be here.

[00:43] Beatrice: So I thought we'd start with talking about why this project happened, because this was a project that was funded by FLI, and yeah, it would be interesting to hear a bit more about what was your reasoning as to why you wanted to see — we call it like an advanced worldbuilding project, meaning that we actually try to find quite senior people within science, tech, and governance, and so on, to give input on building out worldbuilding scenarios around different AI futures. And so it would be interesting to hear your thoughts just on why you thought that was a project that would be worth supporting.

And then also if you have any thoughts on — today we're gonna mostly talk about the tool AI scenario, which is one of the scenarios, but also we did a d/acc scenario. So why those two scenarios, if you have any thoughts on that?

[01:27] Anthony: Great. Yeah. First let me say I thought you guys did a great job on it, and it was so interesting to be a part of it. And I'm really impressed that you got such great thinkers to participate and give you feedback — just excellent work. The motivation for this, I would say, on my and FLI's part, is twofold. I would say one is that, with Foresight, we've long thought that one of the things we really need are just different and positive visions of how the future with AI could be.

And we've been thinking about positive futures in AI for quite a while, because we don't want to just bumble into wherever the incentive structures we happen to be in push us. That's largely what is happening now. I think we all believe that we should have some visions of the future and where we actually want to go — for reasons that we endorse — and direct ourselves in that direction rather than just, okay, let's just do what seems like the thing that we have to do now. And so if you don't build those visions, they won't necessarily happen by themselves. You'll end up on the default trajectory. And so we've always been interested in trying to develop those visions. And we, as far back as the first augmented intelligence summit that we did with worldbuilding — and we've been into worldbuilding, as you guys have, for a while. So worldbuilding was an exciting thing to us, and filling out something that wasn't just a brief little bit, but something really in-depth where a lot of expertise and a lot of effort went into fleshing out how the actual history plays out.

And what the technologies look like, and so on, which you've done such a nice job on. So there's a worldbuilding side. And then I think the other side was there's this ambient sense that there's just a path with AI. We build bigger and bigger AI systems, they get more general, they get more autonomous, they get more everything. And we can either just go down that path and hit AGI and hit superintelligence — and my God, what are we gonna do? Because none of those things — it's very hard to figure out how the world makes sense with superintelligence, or how it's gonna go well for a lot of people with AGI, and how we're gonna control it and how we're gonna have alignment and all those things. But that's just how we're stuck, because either we can go down that path and end up at those things, or we can be Luddites and stop technology, and then somebody else is gonna race ahead and there's no stable way to do that. So the picture is that we can either forge ahead down one path or we can stop. And I think that's just not true.

[03:45] Anthony: And so part of the hope here was to say, okay, why do we think that? And what should we be thinking? Are there actual decisions and design choices that we've implicitly been making so far, and that we will be making in the future, that choose different directions in the technology? Some of those things in these scenarios got chosen for us in the sense that large incidents happened and that kind of directed attention one way or the other, but lots of them were also choices that we made — no, we would prefer to do it this way, that may actually make a lot more sense for a society. And so we're keenly interested in saying — and the truth — which is that there isn't just one direction. There's not just one path we can go down. There are many. And let's look at a few of them that are quite different than the one we're on now, and really get those developed so that people can understand that there are choices and there are alternatives. And that we have a little bit more agency here as a species and a society than just stop or go.

[04:42] Beatrice: Yeah, I think that's a really interesting framing. And obviously we've talked about worldbuilding, I think, literally on this podcast before — probably, yes. Yeah. And I guess also the interesting thing about tool AI and d/acc is that there's sort of, to some extent, like memes that are present if you're in this sphere. And so this may be — people who are interested in these things have probably heard of it — but it's interesting to take them to their extreme almost, or just what does it look like in reality.

[05:11] Anthony: Yeah, because I mean, certainly they're kind of like a philosophy or something, each of them. There's sort of a bunch of thinking behind why we would want to go down that road, and kind of at the abstract level. But it's really nice to make those concrete. What does this actually look like? How would that play out over time? What would the actual technologies be? So it's been really cool to see that in both cases — and inspiring, really, to see it in both cases, in that it's sort of believable that there is this nice future in both directions. And also I think it's good that there are two of them. We're not just saying, this is actually the future we should do instead. We're not going to end up in any future that we imagine right now. That's not how the future works. It's very unpredictable. So the crucial thing is to get more than one option on the table, and see what we like about the different things, and have them be draws for us when we're making decisions — oh we don't just have to pick the party line — and that we can do things now that make…

Insofar as we run into some of the things that happen in these pictures — if we run into large-scale incidents like we did in the d/acc scenario, or if we run into, well, agents are inevitably gonna screw things up a whole bunch, and people are gonna realize that nobody has responsibility exactly, or we don't know who has responsibility for the actions these things take. How are we gonna think about liability?

Running through these scenarios and sort of playing them out lets us think now about the systems and ideas and work that we need to have in place for when some of those things start to happen.

[06:42] Beatrice: Yeah, I think definitely one thing that becomes really clear is also the benefits and trade-offs with the two different versions. Like you can see that there are massive benefits with the d/acc in terms of resilience, but maybe there are some trade-offs in terms of potential for flourishing, or things like that. So maybe we dive into a bit of that for the tool AI one specifically. What do you think about — when you think of a tool AI future? I think maybe we start with what excites you about it, because also arguably you wrote this paper "Keep the Future Human," and that sort of aligns to some extent with a tool AI future. And we use this diagram — the Shield diagram — that you created in this paper to sort of clarify what we mean by tool AI. And yeah, when you think about a tool AI future, why do you think that maybe that's the direction that we should be steering towards?

[07:36] Anthony: Well, I think tools are basically what we've built with our technology most of the time up until now, and have served us really well. Like most of the things that we like about technology — and that have brought huge amounts of wellbeing and value and productivity and safety and all those things — have been tools. And I think if you actually ask most people that are in AI development, "What are you building?" they'll say something like, "We're building cool new tools to do X, Y, and Z." And so they're actually sort of already bought into this paradigm anyway.

I think the thing that is sort of the exception is AGI — as this combination of some of the elements of tools, now this intelligence and this generality, but attached with high levels of autonomy. And that really makes it quite different. And so, as I talk about in that paper, it's that triple combination that makes things very different than tools that we have right now, and also makes a different overall dynamic in the sense that rather than having things that people use — like tools that people use for their own benefit and empowerment, and that are complementary to the skills that they have, like most tools, or extend the skills that you have — they're replacements. That, once you have autonomy, generality, and intelligence — the center of that Shield diagram — right now that's where people are. And so once you can put an AI system there, then you can put them in and take the people out. And that's sort of the picture.

And so I think the upside of tools is the upside that tools have always had: they let us do more powerful things that we want to do and give more power to individuals and groups of people. And so all of the science and technology that we actually want to get out of AI — letting us do things that we currently can't do, letting us do certain things much faster, just giving people more, extending the capabilities that they currently have to ones that they didn't have before — all of those are exciting and positive things that tool AI can give us. The thing that it sort of avoids is replacing people with the AI system. And most people don't want that. They're doing stuff that they want to do and they don't really want to be replaced in doing that. Or they're doing stuff that they have to do because they need to make a living, and they also don't want to be replaced doing that either.

[10:01] Anthony: Most people are doing the things that they're doing for reasons, and don't want to be replaced with something else that does it instead of them. They want something that empowers them. So I think the benefits are quite clear. I think the trade-offs come if you lean out of especially the autonomy and into the other parts. And I think the things that we hope to get out of AGI — systems that can turbocharge science, or let us do higher productivity and just lots more of the things that we're trying to do — I think we can get all of those out of tool AI. Like AGI at some level doesn't do anything that people can't do. That's sort of — if you talk about superintelligence, I think that's a little bit different. But if we talk about AGI — the thing that is human-expert level at all of the things that humans are expert at, and just does it autonomously and much faster and much cheaper — that doesn't really allow us to do necessarily things that we can't do. It allows us to do them more cheaply and replace people doing them. And insofar as it does let us do things we can't do already, tools would allow us to do those things too. So insofar as there are things that we can't do because they're at such a scale and require so much speed that we would need AI to do them, tools can do those things too, I think.

So I think there are trade-offs with AGI, but the primary one would be that it would be able to do things faster than without it, because what the autonomy lets you do is take the human out of the loop, and humans are kind of slow, right? And so if you can have a thousand AGI systems running at fifty times human speed, you're just gonna get a whole lot done. It might be the same stuff that humans would have done eventually, but you can do it a lot faster. And so I think we would get to the same sorts of outcomes, but it could take longer if we had tools instead of AGI. And some people don't like that. So there is a trade-off there between: to what degree do we just race and get things as fast as we can, versus get them in a way that works better for more people and is safer and more controlled and all of those things? With superintelligence, I think it could be different.

[12:08] Anthony:  There may be things that take a very long time to get to great nanotech like next year — you're not gonna get that with tool AI, or with AGI, I don't think. If you're dead set on immortality in the next two years, you're not gonna get that from tool AI or from AGI for that matter. I think it's only going to come from superintelligence. The downside, of course, is that it might lead to disempowerment of humanity and the extinction of the human species and all those things. So there's a huge trade-off there. I think the trade-offs are real, and we should take them seriously, and that goes in both directions. We should see the things that we would give up by not having AGI and superintelligence. But I think those things are sort of less than advertised. And the things that are generally advertised — like AGI is going to help us cure cancer — what we really need are tools for that. The things that I think are slowing us down from getting cancer cures are not the things that AGI is going to give us. They're things that powerful AI tools and lots of other stuff — like investing in the datasets, changing regulatory structures, and all those things — are the things that are standing in our way, not smarts…

[13:27] Beatrice: Well, great. I feel like you've covered a really large chunk there, both in terms of the benefits and the trade-offs. I think one thing that would be interesting to dig into a little bit more is — when we were working on this scenario, one interesting thing was: how stable is tool AI, for example? Even if we were to achieve a tool AI future in the next ten years, do you see it as a stable path? Do you see it as sort of a temporary thing until we understand what AGI is better? What option space do you see if we reach a tool AI future? What happens then?

[14:04] Anthony: Yeah, I think this is really hard to know. I do worry that there's some instability in the sense that people are lazy. And if there's a system that says, "Don't worry, I'll do all of that stuff for you, and you can just trust me, I'm gonna do it well," there is a real temptation to do that. And so this is going to be really a defining question, I think, for humanity: how much we want to retain agency and control. I think there's a level at which you can't have it all.

Like, you can't both delegate things to agents and have them do all the work and also really have meaningful control of what they're doing or agency. And once you delegate things, you've lost a little bit of involvement and understanding and control of how it turns out. And I think, as you imagine very powerful AI technologies, there's a level at which we are gonna have to make a decision. Do we want them to be our tools or do we want us to be their tools? It's going to be a little bit of one or the other. And there are two competing drives. People do want to be in charge. They do want things to go their way. They want the things that they want rather than some random other thing. And that means they want to be in control and they want to have say and agency and freedom and so on. On the other hand, they also want things to be easy, and they want things to get done quickly without a lot of work and all of those things. They also want to spend very little money getting those things done — if they're an employer, or if they're just trying to make something happen, they don't necessarily want to employ a bunch of people and have it be a lot of work. They want to just have it happen. So I think there are pushes in both directions. And honestly, I think it's going to be a sort of central issue of our civilization over the coming decades, how that plays out.

I don't think there are any easy answers there, because there definitely are going to be pushes in both directions.

[16:03] Beatrice: Yeah, it's interesting to just be aware. And in the report we have a little discussion about the different trade-offs and questions. Another thing — when we interviewed people for this scenario, most of them said that they were really excited about a tool AI future, that this was, in terms of "what AI future do you choose now," like a tool AI future — that was my sense at least. But no one thought that we were on this trajectory right now, basically. And many in the interviews said that it was, very roughly — a sort of a little bit hand-wavy, I guess — that the incentive structure isn't there, that's not where we're heading right now. Do you have any ideas on how we actually shift this? And also, just why are all the narratives captured by AGI narratives, for funding and for just our imagination? And how can we actually change that story now?

[17:07] Anthony: Yeah, yeah. It's an unfortunate place that we've gotten ourselves to — I think it didn't have to be this way. I mean, I think there's a historical element here: a lot of this started from people who were thinking about this big picture, long future of AI and superintelligence. And rather than AI — the progress in AI that was happening in reinforcement learning and in Transformers and all of the technologies that have come and that have made these things possible, along with the huge amounts of computation — they could have come into a world where most people were just thinking about AI tools and not a successor species. And that would have played out fairly differently. There was this ideology that was kind of hanging around and informed the founders of some of the companies that were suddenly given access to huge amounts of capital and technological success, and just in the place that was able to make use of the technology that was now becoming possible. And I think part of changing the way that's operating comes down to two things. One is understanding that what most people know — the things that are being promised from these companies primarily to the public and to policymakers — like, it's going to turbocharge our technology, it's going to cure cancer, it's going to allow people to do all these cool things and we're gonna build it into all of our products. All of those are tool things. Those are not things that require AGI to do. And if you ask the companies, "What are you building?" they're like, "We're building these amazing tools." And so at some level they're already on board with what they're saying. The problem is that there's this ideology — and there's also this race where AGI is thought of as this goal, like the shining thing.

That we will attain. And whoever gets to the shining thing first wins this gigantic prize of power and wealth and civilizational dominance and all of these things. I think those things are not true. There isn't a shining thing. We've already seen that even defining what AGI exactly is is quite hard. And there's probably no thing that we're going to get to where everybody says, "Yes, we got it, now we're at AGI." We've seen that even for the AI that we have so far, nobody can even agree on whether it's impressive or not at a given step.

[19:28] Anthony: Right. GPT-5 came out, and some people said, "This is a great advance." Some people said, "This is just totally on trend." And some people said, "My God, this is a total disappointment, scaling is hitting a wall," blah, blah, blah. You get a full spectrum of responses on almost every thing that comes out. I think there will be no point at which everybody just agrees, "Okay, now we have AGI." So AGI is more of a direction in that sense than a place that we're going to get to. But creating a big glowing trophy that gives you unlimited power and wealth and goodness — sure, that is seductive. That is something that you can drive a race around. Whereas "let's build even better technologies and better science" — that's what we actually want, but it doesn't make quite as compelling a narrative. So I think part of what we need are things like what this project does, which is play out: what are we actually talking about in these different scenarios? What are we actually choosing between?

When we talk about one path versus the other — because if one is this nice, compelling, vague but very aspirational package, and the other one is, "Well, let's not do that, let's do something somewhat different," it just doesn't have the sway. So I think part of painting the picture of what they actually look like helps to change the narrative. I think the other thing that is important is for people to understand what it means for them to be in these two different paradigms. So the bad news is that part of the incentive structure for AGI is not curing cancer, or even producing immortality, or even producing the super beings that are going to solve all of our problems. It's replacing people. And I think the real economic driving force behind AGI — I personally believe — is that it is the thing that allows you to not have to re-scramble all the tasks in your company, not have to figure out how to automate certain things and have the right controls over them, but to just say, "No, I'm not going to hire this person, I'm going to hire this AI system instead." Or, "I'm going to lay off 90% of my workforce and just replace them with these AI systems." It might be a little bit rough in the transition, but I'll do that. And if you can — so that's sort of seductive to companies at some level, but it's super seductive to companies providing the AGI, because they are then rather than…

[21:50] Anthony: Six billion — or whatever how many people there are working in the world — you have six billion sets of GPUs. That is a huge part of the world economy that you have captured in your company or whatever. So the economic prize is getting a significant fraction of the world labor market, which is tens of trillions of dollars. That's something that is worth spending hundreds of billions of dollars in data centers and accruing these vast amounts of venture capital and all of those things.

Productivity tools — unfortunately, it's much harder to make that case. You can make it, but it's much harder. So I think unfortunately, what is driving a lot of the activity is toward human replacement. The thing that will change the dynamic of that is humans realizing that that is the dynamic driving most of it — that these companies are not trying to make tools, that they're trying to make replacements. I think people are getting wise to this and are certainly understanding that this is a threat to human labor. And it's a threat to humans in their roles as therapists and teachers and companions and lovers — like, all of it. So I think they are starting to understand that replacement is part of the goal and are reacting to that. But I think until that reaction becomes strong and has enough political and social capital to create a lot of pushback, the thing that is going to be driving this is the economic incentive to create these human-replacing things that can capture huge amounts of economic activity. So I think that's both the problem and part of the solution. The problem is that economic incentive; part of the solution is that social realization and human preferences that don't want to be replaced by GPUs and are going to push back. And whether they can push back soon enough and strongly enough and in the right way to sort of change tracks remains to be seen.

[23:41] Beatrice: Do you have any idea — because you use the term "replacement" — do you have any idea about what you think is the sweet spot in terms of what we want to replace versus not? Because I'm sure some things, and some jobs even, we want to replace, and we want people to be able to have better jobs. We want people to maybe have jobs or things that fill their lives with meaning, but maybe they don't have to work so much. In the scenario, for example, I think the work week is more like twenty to twenty-five hours in a decade. So maybe we're lucky in that way. But yeah, do you have any thoughts on that?

[24:17] Anthony: Yeah, I think we want tasks to be replaced. And insofar as there are jobs that are equivalent to tasks, I think a lot of those are probably going to go away. But humans are not made to do single tasks for the most part. We're very general purpose. We don't enjoy doing single tasks eight hours a day, five days a week, fifty weeks a year. We don't like doing a single task — that is not what makes us feel good. And so I think the things that humans like doing are things that really make use of their capabilities, that are different in general, and that they have some feeling of connection to. So I think — it's impossible to know how this is going to play out — but I think as long as the systems we build aren't deliberately trying to do all of the stuff that humans do in the same way that humans do them, then there's at least room for people to reorganize themselves and retrain and rethink and do shorter work weeks and all of those things. If we explicitly make the goal to make AI systems so that there's nothing humans can do that the AI systems can't do, then obviously what are we going to do? It sort of answers itself. That doesn't mean there isn't going to be a huge amount of disruption if we build lots of tools. We've seen technological disruption before. There will be — you can have autonomous vehicles that are fairly narrow, they just drive, but they're not super general or intelligent. They're tools by this paradigm, but they will put people out of work. People that just drive — there are gonna be fewer of those positions if we have huge adoption of autonomous vehicles. That's gonna happen, and people are gonna be unhappy about that. And there are lots of other things that are gonna be similar. And so the economic disruption is gonna be huge, and we will have to figure out how to deal with that, even in the tool AI paradigm. It's just that I think that's at least possible. It doesn't feel totally hopeless — like it does in the AGI paradigm — for people to have jobs that they actually get compensated for, because they're needed. You can imagine there are other economic systems that we could conceive where people have jobs, but it's hard — it's going to be a different… than what we use the word "job" for. If actually people don't need to pay them to do those things, then we don't actually need them done. So yeah, I think…

[26:54] Anthony: How to find the sweet spot — I think we probably won't, frankly. We'll probably vacillate around and mess it up in a lot of ways. But I think we at least have a chance if we don't go in the full replacement direction. I think the other thing that I find exciting about tool AI — and AI in general — is not just doing the same stuff that we're doing now but more cheaply, but being able to do things that we couldn't do before at all. Some of those things are because we just can't process that information. I think AlphaFold can fold proteins in a way that we couldn't figure out how to do in any other way — we tried a long time, and that was the answer. So I think there are things that our human brains just can't do. They can't fold proteins as well as AlphaFold, and that lets us do things that we couldn't do before. It's not just cheaper and faster. Another thing comes from scale. One of the things that I'm excited about is new ways of doing democracy and social interactions and deliberation. Up until now, it has always been one person talking to one person, or maybe one person talking to a lot of people. If we have a lot of people talking to a lot of people, up until now it's just been a mess — just crowds shouting at each other. But now, with language models, we have the capability to have a thousand people having a conversation with another thousand people in real time. That's something we've never had the possibility of before, and we can now do it. I don't think we're quite doing it, but I suspect that we will. And I'm excited to see if we can make that happen. What does that look like? That enables us to do things that we could never do before — like have a hundred thousand people come to a negotiated agreement on something in a real way that isn't just electing officials that then negotiate for us and don't quite do what we want and blah, blah, blah. So I don't think we should get rid of elected officials — don't get me wrong. But having more tools in the arsenal than what we have now — which is kind of shouting at each other online and talking individually with people and giving our one bit of information every few years to vote for this person versus that person — there are lots more things that are possible that we simply didn't have before and that AI will allow. So I'm very excited about those.

[29:20] Beatrice: Yeah, I think there's definitely a lot to be excited about if we get it right. So another thing that I think is a key question — we started on incentives. And what happens in the scenario is that there's an incident where an autonomous AI system misdiagnoses people, it gets very messy, and this major liability case emerges from AI healthcare. And this leads to the AI liability framework. And this puts us on a trajectory where insurance companies don't want to cover systems that have — we call it the triple intersection, basically based on your Shield diagram, because it's the high autonomy, high generality, and high intelligence. And you're the one who I think suggested that, to the extent that we get on this trajectory, that's probably the most likely way we would get there — because we won't like the liability, insurance companies won't want to cover the triple intersection. Do you think that this is the most likely way that we'll actually get to a tool AI future? And how can we otherwise make sure to get to a tool AI future? How do we get on that trajectory?

[30:39] Anthony: Yeah, yeah. I wish I knew exactly. But I do think — most likely, as you did — I imagine that we get there. How did we get there? I think probably some combination of pushback against replacement: people really understanding that AGI means replacing humans in their entirety and not wanting that. But that alone won't change things, right? Because it's very hard to imagine what a policy solution to that looks like.

So there's a sort of public pressure and backlash that I expect will happen — and whether it happens, and in what way, and what effects it has, we'll see. And then I think the liability side is likely, because up until now, when things go wrong with AI systems, it tends to be the user's fault, right? The primary thing that the AI system does is provide information. It's very much clear that it's the responsibility of the user to vet and check that information before they use it for something. And I think people do feel that way. I don't think if you say, "I'm sorry, I totally screwed that thing up, it's the AI's fault," anybody right now is going to take that as an excuse. It's going to be totally your responsibility — and they'll say, "Well, why did you use a stupid AI system instead of doing it yourself? Why didn't you check the result?" etc. So I think we currently very much have an idea where the responsibility is on the user. And most of what the AI system is providing is information or text or media or something, and taking action is the thing that the user is doing with that information or content. But once we start having AI systems that are agents and autonomous and are taking actions on their own, it's going to be a very different picture. If I set my AI agent to go do something and then it goes and does it and it messes it up — am I to blame? Is the agent to blame? Is the company that provided it to blame? We're going to have to work this out. And I think when the stakes are high for those things, it's gonna be very unclear who's responsible for them. And I think most users are not going to want to be the ones responsible for some AI system screwing up. Nobody's gonna want to use an AI system where, when it screws up, it's your fault.

[32:57] Anthony: And you have nothing — like you have no control over it, it's just off doing its thing, and it's your fault when it screws up. Nobody wants that. So I think it will be problematic as a product, but legally it's also going to be obviously totally problematic. And so the question is how we're going to adjudicate all of that. And I think — in my view — we will adjudicate it in a way, and the legal precedents that we set up and the liability systems that we set up, such that most of that liability — if a system really is not under the user's control, like it's not a controllable system — then the liability ends up on the provider, the company that developed it. It shouldn't be on the AI system. They can't have responsibility. You can't put the AI system in jail or something if it does the wrong thing. So I think responsibility landing on the AI system makes no sense. So it's either on the user or the provider, the developer. And if it's a system the user can't control, I think it should not be on the user. How can you hold somebody responsible for something they can't control? You can hold them responsible for choosing to use something that they can't control, but I think ultimately the liability needs to land on the provider. And if it lands on the provider for an uncontrollable system, we'll have liability that is very high for uncontrollable systems. And that's what can drive you to tools, because they're sort of by definition controllable things. So I think it is plausible that it'll work out that way.

I think it's certain that many more incidents will start to happen when we have agents taking actions, and it will obviously be a problem as to who is responsible for those screw-ups. And we will have court cases adjudicating this. I think what we don't know is how that's going to play out. Will the company somehow skirt responsibility and it's still on the users? Or will the insurance companies just absorb the cost of this — like we do with credit card fraud or cyber attacks — just somebody taking care of the fact that all this goes wrong somehow, and we just take it as a cost of doing business? Will it be something like that, even though the stakes get really high? I just don't know. But I think that is the most likely route: that sort of combination of public pressure and an adjudication of where the responsibility lands for these autonomous systems doing what they're doing.

[35:16] Beatrice: Yeah, no, that's really interesting to hear. So I think one thing that would also be interesting to hear your thoughts on is: how does tool AI compare to comprehensive AI services? Do you have a take on that? Would you say they're kind of the same, or what's the difference?

[35:31] Anthony: I think they're interesting and related, and they play together in interesting ways. I think they both take the standpoint that we should not build a giant monolithic AI system that is a black box that does everything — whether you call that AGI or superintelligence or something. Comprehensive AI services is an offered alternative to that, where it is much more modular: you have particular small-ish AI systems that do particular things and provide particular services. That makes it much closer to a tool-like system because they're more narrow and focused and probably more controllable. So I think there are a lot of overlaps in that sense — it is a different paradigm. I think, however, comprehensive AI services is imagined as being comprehensive, and that includes having things that are very autonomous. And so I think comprehensive AI services is, as I understand it, conceived as a sort of another version of AGI superintelligence, but one that isn't a monolithic black box — it's this kind of federated thing that is working together to build something that functionally does all of those same things. So in particular it could still have generality and autonomy and intelligence all bound into the system. But there'd be lots of advantages in having those be modular; it would still be able to do all of those things.

So I think, if I had to choose between a giant monolithic black box AGI superintelligence system and comprehensive AI services, I would totally pick comprehensive AI services, because it has lots of advantages in terms of safety and human intervention. There's no way for a human to really intervene in a giant black box superintelligence system. You can at least hope to intervene in a comprehensive AI services system, because there are connections between the different services and modules where a human can fit and oversight can fit and so on. And it doesn't have to operate at some fixed speed — you can have regulators in it that have approval steps and human intervention and all those things. So it's much more human-friendly and control-friendly than a big monolithic superintelligent system. I think it's preferable. I think it does, however, also still have the risks of…

[37:58] Anthony: You can imagine a comprehensive AI services system that you just treat as a black box anyway. Even though it's broken into pieces, it's got a super agentic autonomous module that runs all the other little modules to do all the individual things that are more like tools. And you just say, "Okay, autonomous module, go do my thing for me," and it goes and does it. And in that sense, it's very much like an AGI system. So it allows you to, but doesn't force you to, have humans involved — the way that tools do. Tools require a person to operate them and an AGI system doesn't. And that's the crucial thing. So I think comprehensive AI services would be better because it would allow humans to intervene, but it wouldn't require them to. And so there would still be the danger that you overdelegate and just let the services systems be more like AGI and kind of run away with things, because they could. It also shows a sort of danger of the tool AI system, which is that it's not so hard — if you have a very capable, general intelligent tool — it may not be that hard to turn it into an agent or an autonomous system, if you just add a little autonomy module that says, "Okay, now what do I do next? Go do that thing. Now what do I do next? Go do that thing." A planning module and a doing module — then you've got an autonomous system. So it's not so easy to build things that are very capable but still very tool-ish. And I think that is a real design challenge. What does it mean to build very powerful but tool-like systems? It won't happen by itself — that is something that you actually have to design in. And I think that's a really interesting scientific and engineering question: how can we do that?

[39:50] Beatrice: Yeah, that was really interesting to hear your take on the differences. Last question. Going back to what excites you about tool AI — it would be interesting if you have any ideas for specific tools that are maybe shovel-ready, ideally something more concrete than maybe AI for healthcare. Do you have any concrete and exciting ideas that we can get to work on right now in terms of tool AI?

[40:13] Anthony: Yeah, the ones that excite me most — I would say — are the coordination one, which I discussed before: deliberations in a way that we haven't been able to do before. Because I feel that one of the crucial things we're lacking right now is an ability for large-scale human preferences to be compared and aggregated and sort of worked over and deliberated on and put into action in a way that is just much more responsive and fast than the systems that we have right now. I think most people don't feel like our political and other institutions are really channeling their preferences into action, probably. I sure don't. And I think most people feel quite frustrated with that. And I think new ways of doing new social and technical constructs can help with that. So I'm very excited about that. And that's totally tool-based. And I think it can be totally done with the level of AI that we have now. Another thing that I am quite excited about — and can be done with the AI that we have now, it's just a question of building it — you have it in the report: what I call the epistemic stack, the ability to say, here's a piece of information or data or understanding, like a paragraph in a newspaper article. Where does it come from? In a scientific paper, if you say some statement, it either follows from something in that same paper, or often it follows from some other paper — you have a citation to some earlier work that it's calling on. And you look up that paper and it has citations to other papers. And A, this could be done better in science in various ways that can leverage AI, but B, this could be done for everything. There is no reason why, in reading a newspaper article about something, you shouldn't be able to trace back where that quote came from or where that piece of information came from. How do I know whether to trust this thing? It's some controversial topic like what's happening in the Middle East — did this thing happen or did it not happen? People are saying totally different things. How do I adjudicate that? I should be able to follow the citations all the way back to either somebody thought something, or here's a piece of data that is registered on the blockchain from some device signed by its secure enclave. We should be able to have this stack that we can follow all the way from the high level back down to the raw ingredients, and then be able to figure out: how much do I trust each of those steps? If this came from this person, why should I trust that person? Have they been right about things in the past? Do they have good models of the world, etc.? If it's an information source, how accurate has that thing been in the past? Has it been reliable? Has it made stuff up, et cetera? So we can have a much, much better system for trust and understanding of the world than we have now, which is both hacked together and getting made manifestly worse by AI at the moment. And there's no reason AI should be making our ways of understanding and processing information about the world worse. It is, but there's no reason for that. It's a tool that we should be able to use to make those better. And so I think there's just a ton of low-hanging fruit there that almost nobody is plucking, and we should get into that right away.

[43:30] Beatrice: Yeah, and the epistemic stack was a really good one. And yeah, we hear a lot about how technology is making a mess of our information systems right now, but that's definitely a way that it could help. That's it for now. Thank you so much, Anthony. I think this was really great to dig into a lot of the uncertainties and a lot of the more interesting nitty-gritty of the tool AI future. So thank you so much.

[43:53] Anthony: And thank you so much for all your hard work on this project. It's inspiring and exciting to see it out there, and I hope it really has a big impact.

Read

RECOMMENDED READING

AI Pathways and the Tool AI vs. AGI debate

  • AI Pathways hub: our worldbuilding project, with two deeply researched positive scenarios of our future with AI:
    • Tool AI: A world shaped by advanced but controllable AI systems, powerful tools built to extend human capabilities without pursuing too much general or agentic intelligence. 
    • d/acc: A world shaped by the concept of d/acc: decentralized, democratic, defensive, and differential acceleration of technological progress. 
  • Keep the Future Human, by Anthony Aguirre: the essay discussed throughout the episode, arguing for keeping AI systems as tools rather than autonomous, general agents. Includes the "shield" diagram framing intelligence, generality, and autonomy as the three factors that make a system risky when combined.
  • Control Inversion, by Anthony Aguirre: Anthony's own follow-up paper arguing that as AI becomes more intelligent, general, and autonomous it stops bestowing power like a tool and starts absorbing it.
  • AI vs Cancer, by Emilia Javorsky: Makes the case that the real barriers to curing cancer are systemic rather than a lack of raw intelligence, and that focused AI tools already help today while the "superintelligence will cure cancer" narrative overshadows them.
  • The Pro-Human AI Declaration: A declaration backed by a broad coalition of organizations and individuals, echoing the episode's case for keeping humans in charge, with five principles covering human oversight and override, corporate liability for AI harms, and protecting human work and relationships from replacement.
  • Protect What's Human: The Future of Life Institute's public campaign against AI built to replace people, reflecting the replacement-versus-empowerment tension Anthony describes as the central economic driver behind AGI.
  • Reframing Superintelligence: Comprehensive AI Services as General Intelligence, by Eric Drexler: the Future of Humanity Institute technical report that reframes AGI as a federation of modular, narrower AI services rather than one monolithic system.
  • My techno-optimism, by Vitalik Buterin: the essay that introduced d/acc (decentralized, democratic, defensive/differential acceleration), the second scenario in the AI Pathways project referenced alongside tool AI.

Organizations

  • Future of Life Institute: the organization that funded the AI Pathways worldbuilding project, led by Anthony Aguirre as executive director.

To learn more about key concepts mentioned in the conversation

  • Artificial general intelligence (AGI): AI with human-like, cross-domain intelligence and the ability to teach itself. Explainer by AWS.
  • Superintelligence: intelligence that would exceed human capability across the board. Explainer via The Conversation.
  • Transformers: the neural network architecture behind ChatGPT and similar systems, mentioned as one of the technical breakthroughs that made today's AI progress possible. Overview on Wikipedia.
  • Reinforcement learning: a training method where AI systems learn by trial and reward, referenced by Anthony as one of the technologies that fed into the current AI boom. Explainer by AWS.
  • AI agent liability: the unresolved legal question of who is responsible when an autonomous AI agent causes harm. Overview by Brownstein Hyatt Farber Schreck.
  • AlphaFold: Google DeepMind's protein structure prediction system, cited by Anthony as an example of AI doing something human brains simply can't do on their own.
  • Worldbuilding as a foresight method: the practice of building out detailed, internally consistent future scenarios. Overview in the Journal of Futures Studies.
  • AI scaling laws (and whether they're hitting a wall): the debate over whether making AI models bigger keeps producing proportional gains, referenced when Anthony discusses the mixed reactions to GPT-5. Explainer by Platformer.
  • Compute governance: regulating AI by controlling access to computing hardware, one of the safeguards proposed in "Keep the Future Human." Explainer by AISafety.info.
  • The epistemic stack: infrastructure for tracing a claim back to its sources and evaluating how much to trust each step. Overview by Oliver Sourbut and Ben Goldhaber.