
/
/
58 min
Stochastic Parrots
Emily M. Bender
Professor of Computational Linguistics
About this episode
Few people understand language quite like today’s guest. Emily M. Bender is a Professor of Computational Linguistics at the University of Washington.
She’s also the person who coined the term ‘Stochastic Parrots’ in the now infamous paper that saw Google fire its Ethical AI lead; and the co-author of The AI Con, a book which is unflinching in its criticism of both AI boosterism and doomerism (in fact she objects to the term AI).
I sat down with Emily to discuss topics such as:
How language really works
What LLMs are actually doing
What Stochastic Parrots really means
Why doomers and boosters are two sides of the same, misleading coin
Full transcript
Humans in the Loop
Emily M. Bender (00:00)
That’s not what democracy means. It’s not a few people in power share a bunch of stuff with everybody else. It’s that the power is actually shared.
Seb (00:07)
There’s a topic about control here. humans inconvenience me because they have their own opinions and They didn’t want to come to the office post COVID
Emily M. Bender (00:15)
the more people understand this, the less likely they are to fall for the illusion that chat GPT thinks things, knows things, reasons about things, the booster narrative is AI is a thing, it’s inevitable, and it’s going to solve all of our problems. And the doomers say AI is a thing, it’s inevitable, and it’s going to kill us all
Meet the computational linguist
Seb (00:43)
Welcome back to Humans in the Loop. Today, we’re going to be stress testing some of the biggest claims made around AI. And I’m excited to be joined by Emily M. Bender Professor of Computational Linguistics at the University of Washington, the co-author of the AICon and co-host of the Mystery AI Hype Theatre 3000 podcast. Quite a mouthful. Emily, welcome to the show.
Emily M. Bender (01:04)
Thank you so much. Yeah, we did not come up with a short title. It might still be pithy, but it’s not short for the, for the podcast.
Seb (01:11)
Yeah, you weren’t thinking of fellow podcast hosts when you came
Emily M. Bender (01:14)
No.
Seb (01:15)
up with that one. Emily, there is so much noise, of course, in this space right now. So I find it more important than ever to understand, whenever we’re talking about anyone’s perspective, to understand what are their expertise and credentials that they bring to this conversation and what is their skin in the game. when they’re discussing this. So maybe let’s start with what does a professor of computational linguistics actually do?
Emily M. Bender (01:44)
Yeah, no, I think that’s really appropriate. And I like to actually introduce myself by explaining what linguistics is and what computational linguistics is. So linguistics is the study of how language works and how we work with language. And that encompasses quite a bit. That’s sort of the shortest description that I’ve come up with. Computational linguistics is using computers in the mix, either just as a tool as you are studying how language works and how we work with language or with the end in mind of building language technology. So that is where I come from. Turns out it’s a hugely important field at this point in time because the stuff that’s getting sold as artificial intelligence, there’s many different kinds of technology in there, but the one that’s driving everyone crazy right now, I think is a fair thing to say, is a kind of language technology. And so understanding how large language models specifically, and then chatbots built on top of them are built, how they work is really important. but then also understanding something about how we react to language that is coming out of one of these machines is also relevant. So that is why I think linguistics is kind of having a moment.
Seb (02:54)
Yeah. Yeah. And I guess the computational part of that title, as you say, there’s this new wave of what you say is being sold as AI. How much were large language models and these kinds of things, before they were cool, before like chat GPT came out, before the world went crazy, how much were they a part of what you were? doing in your kind computational study around linguistics.
Emily M. Bender (03:22)
Yeah, so computational linguistics as a field has long had a bifurcation between statistical and symbolic approaches. And originally, my research was fully on the symbolic side. So I do something called grammar engineering, which is basically automatic sentence diagramming, writing down the rules of a grammar so that you can go from a surface string to a semantic representation or vice versa. And I’m very interested in looking at that across languages. So that’s the kind of work that I was doing. but language models as a kind of technology are super old. one start point that you can pick is the work of Claude Shannon in the 1940s. and language modeling as a technology is a really important component of things like spell check, automatic transcription, machine translation. And so in my teaching, because I teach computational linguistics, I have been teaching about language models for a long time. And, they were, cool within computational linguistics before OpenAI made chat GPT everybody’s problem. And in the late 2010s, there was a lot of work that was basically using improvements to language modeling technology that were based on so-called neural architectures, but basically representing words, not in terms of the letters that make them up, but in terms of what other words they co-occur with is hugely powerful in terms of improving all kinds of language technologies. So the field of computational linguistics through the second half of the 2010s was kind of all about that. And from about 2019, maybe a bit earlier, there was a lot of people who were arguing that these language models were actually understanding the text. And that’s sort of where I entered the conversation because as a linguist, I can tell you they are not. And so I was having that argument sort of field internally, and then it became a little bit bigger. with like remember Blake Lemoine, the Google engineer who thought that Lambda had somehow become sentient. So it became this sort of like broader news story, but just
Seb (05:23)
Mm-hmm.
Emily M. Bender (05:24)
still sort of a special interest thing. And then with the release of ChatGPT, it was everywhere. And let me tell you, so ChatGPT comes out November 30th, 2022. The first half of 2023, so January through June, I counted five work days with no media contact.
Seb (05:42)
I imagine that wasn’t the case beforehand.
Emily M. Bender (05:44)
No, I mean, I’ve been talking to the media, but nothing like that.
The Stochastic Parrots origin story
Seb (05:47)
where in this timeline you co-authored, the paper that coined this term stochastic parrots that
Emily M. Bender (05:56)
Mm-hmm, mm-hmm.
Seb (05:58)
hit the headline, shall we say, at least in part because I think one of your co-authors there ended up being fired by Google and there
Emily M. Bender (06:07)
Mm-hmm, mm-hmm.
Seb (06:08)
was sort of fallout of that. So I guess how did you end up being a co-author of that stochastic parrot’s paper?
Emily M. Bender (06:16)
Yeah, yeah. So that paper was written in September, October of 2020. And what happened was Dr. Timnit Gibru, who was the erstwhile Google researcher who you mentioned there, had contacted me via direct message on Twitter asking if I knew of any papers sort of just summarizing the problems with trying to make language models bigger and bigger and bigger. And the reason she was asking is that at that point, OpenAI had released GPT-3 and it was the biggest one and the people around her at Google were like, why don’t we have the biggest one? Let’s make it bigger. And she was a research scientist at Google and co-lead of their ethical AI team. So it was literally her job to like research what could go wrong and write papers about it. So she was like, is there anything about this? And I said, no, I don’t know of any such thing. And then I said, but a paper ought to cover and listed off about six things. And the next day I said, hey, this looks like a paper outline. Do you want to write the paper? And everybody’s busy, but she said, well, let’s see. And she looped in four other co-authors at Google, and I brought in a PhD student. And we put together this paper, which is a survey paper, that means that we weren’t doing additional empirical work. We were basically gathering stuff from the literature that the seven of us had collectively already read and submitted it to a conference called Fact in early October. So this is in terms of the AI craziness that’s happening, this is really early. We’re writing this in late 2020. It was eventually published at the conference in 2021, but of course everyone became aware of it in December of 2020 when Google fired Dr. Gibru over the paper. And then later also Dr. Margaret Mitchell, who was one of the co-authors and Dr. Gibru’s co-lead of the ethical AI team, a few months later was also fired.
Seb (08:07)
I’m sure some people listening are familiar with the term stochastic parrot, which has become known and potentially that paper as well. But for those who don’t, what is it that that paper was really saying about large language models and specifically the term stochastic parrot, which seems to have stuck around? What
Emily M. Bender (08:29)
Yeah.
Seb (08:29)
is it you meant when you use that term?
Emily M. Bender (08:32)
Yeah. So the first thing I want to say is that the paper was not mostly about using language models to create synthetic text, which is what Stochastic Parrots was an attempt to make vivid. We were talking about the environmental impacts focused on energy and carbon impact. We weren’t tuned into the water issues, but the sort of environmental impacts. We were talking about the way that these systems will encode biases. And then if you embed them in some larger systems, you’re basically reproducing biases. And we were talking about impacts on the research field. So if everybody’s placing all of their eggs in this one basket, and it is furthermore a basket that is very expensive to go play in, and you can’t get published if you’re doing something else, then we’re sort of cutting people out. Also talking in detail about how the data curation practices end up sort of necessarily terrible once you’re at a certain scale, because it’s just not possible to do good data set curation and documentation. really the bulk of the paper. And then we had this section, because OpenAI was using GPT-3 to synthesize text, but nobody else was really excited about it. We thought about that a little bit. And there’s some real clear problems there that have unfortunately only been borne out. So the idea with the phrase stochastic parrot was to make vivid what’s happening if you use a language model to create synthetic text. Because What comes out as these models got quite big and GPT-3 was on the cusp of this looks really plausible. It looks fluent. looks coherent. And because of how we interpret language, we are primed to fall for that. It looks like some thinking entity had a thought and figured out how to express it. That’s not what’s going on. What’s going on is that these are systems that are modeling the distribution of bits of words in their input training data. And then being used to repeatedly answer the question, what’s the likely next word? And so stochastic parrot, no shade on actual parrots who are lovely creatures. And for all I know, do have internal lives. think instead of the English verb to parrot, which means to repeat without understanding. So this is stuff coming back out of the training data with no understanding on the part of the system, but not verbatim. So stochastic means randomly according to the probability distribution.
Why “artificial intelligence” is the wrong name
Seb (10:57)
From Stochastic Parrots, the book you co-authored is titled The AI Con and the podcast, as we said, Mystery AI Hype Theater 3000. I guess there’s no prizes for people guessing that you hold some skepticism around AI. I know you even hold some skepticism around the use of that terminology in relation to this technology. But What is your current stance on what people currently talk of as AI and how did you come to that view of this technology?
Emily M. Bender (11:33)
Yeah, so I think the very first thing is that the phrase is problematic on many levels. It is anthropomorphizing to talk about these systems as if they have intelligence. Dig one layer deeper, the whole notion of intelligence is just like a racist cesspool. We don’t need to go there. The idea that you can line people up according to one particular property and that property is sort of this key idea of general intelligence comes to us straight from eugenics and race science. Anytime someone goes in and tries to define AGI. So there was a paper out of Microsoft, not a paper, was pre-print, it was never published, called Sparks of AGI. And their attempt at defining it, if you went through, was based on a Wall Street Journal op-ed written by some horrible race science psychologists, sort of claiming that this is the mainstream view of intelligence and psychology. There was another recent thing that was, I think, collaboration across a bunch of sites that didn’t even really hold a paper shape, but it was presented as if it were a paper. Again, trying to define AGI. And again, if you click through, like you quickly find that what they’re leaning on is other people doing race science. So intelligence is problematic, even if it weren’t artificial intelligence is anthropomorphizing. And even if all of that weren’t problems, it’s also lumping together disparate technologies. There’s not one thing out there that is getting ever better at you know, writing sonnets about getting a peanut butter and jelly sandwich out of the VCR and predicting the structure of folded proteins and picking out the faces in the picture so that your camera can focus on that. And, you know, on and on and on. These are separate technologies. So I think it’s really important to disaggregate and get specific about what we’re automating. So that was just sort of the first part about the phrase. My question to you is when you’re asking me about my stance on these technologies, which ones do have in mind?
Seb (13:31)
Yeah, I mean, that’s a great question. if we focus on the most visible commonly used examples here, if we focus on the chat GPT type of example, if we focus on some of the more common AI agents as they’re being called. I appreciate there’s nuance here and as you say, maybe different stances according to different technologies. But yeah, maybe let’s start there and see where we get
Emily M. Bender (14:03)
Okay.
Seb (14:03)
to.
The illusion of a mind behind the text
Emily M. Bender (14:03)
Yeah, so let’s start with chat GPT and chat bots and large language models. There’s a very important thing to know about how we interpret language, which my hope is, and this is why I’m explaining this over and over again, is that the more people understand this, the less likely they are to fall for the illusion that chat GPT thinks things, knows things, reasons about things, et cetera. So you might think that when you are listening and understanding something that somebody said, What happened is they had an idea, they packed it into words, they sent the words across to you somehow, and you just unpacked the idea from the words. In fact, what happens is much more complicated and way cooler than that. When we, and this is from psycholinguistics and pragmatics, which are subfields of linguistics, when we understand language, what’s happening is we are keeping in mind everything we know or believe about the speaker’s state of mind. So what they know and believe, what we have in common ground with them and… what they know and believe about their intended audience, which sometimes we understand to be somebody else other than us. And then against that background, we use the words that they said as a particularly rich clue to the idea they’re trying to communicate. So we ask ourselves, what must they have been trying to convey by choosing those words and in that order? This is cool. It’s fascinating. Like pragmatics and psycholinguistics are really cool areas of study. Unfortunately, we can’t turn it off. And so when we see some text that comes out of ChatGPT or Claude or Gemini or whatever, in order to interpret it, we have to imagine a mind behind the text. We’re doing it instinctively, reflexively, and there’s this additional step we have to take now that we have these machines in our information ecosystem to remind ourselves, no, that’s not where those words came from. Yes, I can make sense of them by imagining somebody was saying them, but in fact, what they are is someone Set up a system to repeatedly answer the question, what’s a likely next word? So the only evidence we have that something like chat GPT, the only evidence in scare quotes is intelligent in scare quotes or on its way to some kind of general purpose reasoning system, let’s say, is the fact that it outputs this plausible sounding, coherent seeming text.
Seb (16:21)
And I guess when you maybe when you add to that the sometimes quite sycophantic top and tails to that, you’re like, great job, or, know, whatever, whatever,
Emily M. Bender (16:32)
Yeah.
Seb (16:32)
like positive reinforcement, you know, it does, it does give you, you know, the these, this very realistic, on some level illusion of something being on the other end of the conversation. And so yeah, I guess you can see how people fall into that.
Emily M. Bender (16:47)
And depending on what kind of a mood you’re in, that very affirming positive feedback might make it so that you want to believe even more that there’s something there holding those feelings. Or you might find it completely off-putting. Like, I don’t think that’s going to have a universal reaction. do you know why it is sycophantic like that? How did it come to be that way? Right? Because if you just trained your language model on like social media text, you absolutely would not get sick of fancy. Right? That’s the, that’s not going to be the likely next words. So there’s a additional training step called reinforcement learning from human feedback where many, many data workers, some who are, you know, trying to eke out a living doing this and paid very poorly, but also anytime you gave a thumbs up on a chat GPT response, you were also giving the human feedback for reinforcement learning for human feedback. And that allows open AI to shape the probabilities inside their product so that likely next word is not just likely according to the original pre-training data, but likely to be part of a sequence that would get positive feedback from a rater. And that’s how you train in sycophancy.
Seb (17:54)
Interesting. so, you if I, if I hear you correctly here, ultimately, you know, the large language model, as you said, is, is trained. It’s a probabilistic machine that’s constantly asking the question of, what’s the most, most likely next word in the sequence, that there are. reasons or mechanisms through which we as humans interpret language that mean we’re of primed to fall for that as being a kind of thinking being on the other side, even if we know that the inner workings are not quote unquote intelligent.
Emily M. Bender (18:36)
Not what we are imagining them to be. Yeah.
Seb (18:38)
Yeah. I guess one of the things I have observed, know, when we use this word intelligence, which you’ve already said has its own problems associated, but there’s been this very sprawling debates about, you know, the nature of consciousness, of intelligence, of… Basically, some people saying, well, so what if the machine doesn’t truly understand in the way that our brains understand and compute in the way our brains compute? The fact that the output is of the quality that people perceive it to be, is that not intelligent? don’t know you’ve come across that argument and if so, like what do you make of it?
AI scribes and the doctor’s office problem
Emily M. Bender (19:29)
Yeah, I mean, think I’ve seen forms of that. and I think there’s sort of two things going on there is like, why do you feel a need to establish this thing as intelligent? Like what’s the, what’s the purpose of that? What follows from that? but then also, how do you know it’s actually any good? Right? This is one of those things where, it’s something that looks very polished. We tend to think of as more likely to be correct. and. Even if, one, I’ve done a bunch of writing starting in, well, the first publication was 22, but I started on the work in 21 about why we do not want to use the chat interface as an information access system. And there’s many reasons for that, but one of them is even if it’s right most of the time, that’s actually arguably worse than it being less reliable because something that’s going to be right 95 % of the time. How are you, is the user going to know when you’ve hit the 5 % or it’s wrong? You’re specifically using it for information access, meaning you don’t know. There’s another thing that’s happening a lot recently is people are pushing so-called AI scribes into the doctor’s office. So the idea is that it records everything said during the patient encounter and then produces a first draft chart note, summarizing that. And there are horror stories all over the place about incorrect things ending up in chart notes. Because the selling point here, the value proposition for the physicians is you don’t have to spend all that time charting. All you have to do is check to make sure it’s correct. But that kind of checking is actually really hard work. And to do it thoroughly ought to take as much time as just writing the chart note in the first place. So that doesn’t happen. So.
Seb (21:15)
Have you, you, have you watched the TV series, the pit on, on HBO?
Emily M. Bender (21:18)
No, but I hear this was there was a plot line.
Seb (21:20)
Yeah. I would try it. It’s one of my favorite series right now. And there was exactly this plot line, you know, doctor who, you know, they, they start trialing, you know, AI scribe. She clearly doesn’t check the thing. It gets sent off to another department and then,
Emily M. Bender (21:35)
Yeah.
Seb (21:35)
into the scene comes running a furious surgeon from upstairs because she’s got. none of the right information or some key information that’s just very clearly not right. So yeah, I think you can totally see that. you know, I have seen it and I know many people report it, this notion of work slop, this technology will produce something that is kind of the illusion of good work pretty quickly. And then people just start pushing that around an organization. You your boss asks you to do a presentation on something rather than do the thinking that makes the presentation good. You kind of chuck it into, you know, chat GPT or Claude. It comes back with something. You don’t really check it. You send it off. And, and, and, and thereby, so like push. push the burden really onto somebody else to figure out like, this, is this actually good work?
Emily M. Bender (22:27)
And a lot of what’s happening sort of surprisingly over and over again in different sectors is mistaking the words that we say, the slides that we put together, the diagrams that we draw for the work that we’re actually doing. Because in many cases, that is the most immediately observable part of the work. But if you think about what it is to be a doctor, a lawyer, a teacher, a therapist, a fitness coach, any of these roles, putting out words that sound like someone in that role would say is not actually a replacement for that work. But many of these cases are cases where we have concentrations of finances, right? We put a, not enough, we put a lot of money into our healthcare system and not enough, we put a lot of money into our education system. And so the people selling these products would love to tap into that. And the people running the systems are in sort of a continual state of scarcity and austerity. And so if Microsoft, OpenAI, Google are selling things at a loss, by the way. Nobody’s making money on their synthetic token extruding machines right now. If they are selling it cheaply for now, that looks like a quick fix to not enough time for the people to do the job because it’s too expensive to hire people, which is really dangerous, right? Because what happens if and when OpenAI goes under or raises the prices? And if you have reshaped systems around these already bad substitutes and lost your workforce, things are going to be even worse.
Seb (24:00)
Yeah. Yeah. And I think you could very reasonably argue that you’re sort of papering over the cracks of a system that is broken in other ways and kind of saying, right, here comes this like silver bullet. I guess that is very much seems to be the Silicon Valley booster narrative of just like everything gets solved with AI. We move to this sort of new new phase of humanity in which everything’s abundant. I don’t know, we’re all getting universal, not even universal basic income anymore. It’s now universal, you know, good income or whatever the right term is.
Emily M. Bender (24:40)
A fully automated luxury communism is the thing.
Unpacking The AI Con
Seb (24:45)
That’s a good term. I read the AICon and very much enjoyed it. And as I was reading it, I tried to, for myself, summarize the issues you highlight as I heard them. So perhaps I can kind of share this list back with you and get your reaction and see what I missed. And we can build on some of these topics. So the list as I’ve got it here. So the first is the claim of democratization. you know, it’s a big part of the narrative or everyone has access to X, Y, know, super intelligence. you sort of label this, a false promise for a whole host of reasons. You also point to a lot of the kind of degraded jobs and negative labor market effects. so there’s, there’s a, you know, a big topic there. you talk about the social issues that are built into the tech. so very obvious, example might be racist facial recognition technologies
Emily M. Bender (25:45)
Mm-hmm.
Seb (25:47)
point to the huge copyright theft. people argue both sides of that and say, yes, it is, no, it isn’t. the use of huge volume of materials to train these models, the environmental impact in all its forms, the proliferation of slop in all its forms, the
Emily M. Bender (26:07)
Mm-hmm.
Seb (26:07)
anthropomorphization, which is a hard word to say, of machines. And then there seems to be maybe less explicitly, but I guess that to me that felt like a tone or understandable tone of mistrust of some of the people ultimately who are building these technologies or leading in these spaces. And the last one I have on this list here, which I’d be interested to talk more about is this idea that kind of the so-called do-mers and the boosters are actually sort of equally misleading and equally almost two sides of the same narrative coin here.
Emily M. Bender (26:46)
Mm-hmm. Mm-hmm.
Seb (26:47)
So anyway, there’s my long list for you. I just want to sort of put that to you for a second and sense check if that feels like a fair summation of some of these different topics and maybe there’s some stuff we want to build on that.
Emily M. Bender (26:58)
Yeah. I mean, I wouldn’t want to say that it gets a hundred percent of what’s in the book, but I think that that is a good sort of sample of the threads that are in there. I wanted to react to the word democratizing because democracy is not shared access. Democracy is shared governance. And when you say, you know, distrust of the people building this. Yeah. I absolutely do not trust anybody who builds their tech by polluting the environment, exploiting workers. stealing other people’s work to be working in the good of humanity. Like they clearly are not in a position where they see the entire rest of the world as people. And so I distrust that. And when those same people say, see, we are democratizing this by making it accessible to everybody, like they’re lying. That’s not what democracy means. It’s not a few people in power share a bunch of stuff with everybody else. It’s that the power is actually shared. And then secondly, I don’t believe they’re everybody because they don’t really see the whole world as fully human.
Seb (28:03)
Yeah, yeah. This narrative around, doomers and boosters or whatever terminology you want to do, doomers, zoomers, there’s all sorts of words flying around. But basically the people, if we want to put it bluntly, the people who kind of do the podcast circuit, the talking circuit, writing articles, essentially saying AI is going to steal all of our jobs. It’s going to tank the economy. It’s going to… Fill in the blank here, huge, world impact in a negative sense. And then you of course have the, those who are huge kind of boosts of technology who say it’s the next industrial revolution. It’s going to change the world for the better. We need to lean into this a hundred percent. need to deregulate. We need to, you know, full speed ahead. Yeah. And I found it an interesting perspective that to see these as almost two sides of the same. So yeah, be interested, just get your take
Emily M. Bender (29:01)
Yeah.
Seb (29:01)
on that.
Emily M. Bender (29:02)
So the, I think it’s easiest to see the fact that they really are two sides of the same coin. If you see that the booster narrative is AI is a thing, it’s inevitable, it’s imminent, and it’s going to solve all of our problems. And the doomers say AI is a thing, it’s inevitable, it’s imminent, and it’s going to kill us all or otherwise destroy everything. That’s the same story with just to choose your own adventure twist at the end. And it’s all predicated on the idea that if we just make the models big enough, then the racist piles of linear algebra are gonna combust into consciousness and have their own desires and be able to self-improve. It’s all speculative fiction. It’s all nonsense. And unfortunately, it’s extremely well-funded nonsense and it is starting to get more and more hearing in the halls of power. So Bernie Sanders has sort of become a Lazare Kowalski alkylite, alkylite around these things. And he’s, you know. talking about how it’s possibly gonna, hosted this thing with Max Tegmark and I forget the other folks on there, but basically exploring the Doomer narrative and saying, Sanders is up there saying, why don’t we have more attention being paid to this in Congress? It’s like, because it’s nonsense, right? Like there are real harms. We do need regulation. We do need to apply existing regulation, but we don’t need to be spending time worrying about. this apocalyptic scenario, there’s plenty of damage being done, but that damage isn’t being done by AI. It’s being done by people who are making horrible choices, either very unsafe choices or very much stock price driven choices. Laying off people looks good for the stock price. Saying that you’re doing it because of AI looks doubly good for the stock price, so do that.
Has anything changed your mind?
Seb (30:47)
one of the things here, you know, almost to, to, suppose, lay my own cards on the table here for a minute, I, perhaps like many people, I have these real swings in my opinion, truthfully on, on, on the topic in the sense that, you know, I am of the tech industry in terms of my career and newsflash People in tech lie all of the time. And we have very public examples of this. Elon Musk has been promising that Tesla is going to be fully self-driving for years. you could draw out a timeline of like, it’s six months away. it’s a year away. it’s
Emily M. Bender (31:29)
Mm-hmm, mm-hmm.
Seb (31:30)
going to happen. And yet every time he says it, it seems to be sort of swallowed hook line and sinker with people who are blind to the fact that he said it. 100 times before on different timescales. And it’s always just the next year, the next timeframe. So I am fully bought into the belief and I have many experiences of my own firsthand where tech companies, the system incentivizes a level of, if you want to put it charitably, of bravado and bluster and fake it till you make it. If you want to be less charitable about it, you know, outright lying and, malign tactics. And at the same time, the actual capabilities as, as I sort of see them and interact with them do shift. And it makes me kind of try and reconsider, okay, what do I, what do I think? And, and the difficult thing for me, and as I’m sure is true of lots of other people is sort of separating what is it I want to be true versus. What is it I actually see in front of me and believe to be true about this thing? And so I guess my question in all of this is, has your opinion, I suppose, shifted at all over the timeframe that you have been actively writing, talking about podcasting in this space? Like has anything come along that’s made you rethink any of the the skeptic view that you you hold about the you either the zoomer or the doomer scenarios.
Emily M. Bender (33:03)
No. No, Definitely not. There’s been more and more hot air, more and more noise. Every time there’s this claim of emergent capabilities, it is paired with an utter lack of transparency into how the systems were built. And so when you say you need to differentiate between what you want to be true and what is true, what you’re reaching for there is a scientific method. How do you construct a rigorous experiment to tell how something is working? And one of the things that is lacking in the whole AI discourse, but has been lacking for a while actually in machine learning as a research field is really good evaluation practice. And I’ve also written about this. have a paper, lead author is Deb Rajee, who’s an amazing author and researcher. And the title is AI and the everything in the whole wide world benchmark. And this is from 2021. And it’s based on the title alludes to a a children’s book based on Sesame Street from the 1970s called Grover and the Everything in the Whole Wide World Museum. And in the book, Grover goes into this museum, advertises the Everything in the Whole Wide World Museum, and he sees a room full of things that are soft and fuzzy and a room full of things that are very light and a room full of things that are very heavy and on and on on, different categories. And he says, hmm. I’ve seen many things in this museum, but I haven’t seen everything in the whole wide world yet. And then he comes to a door that says everything else, which of course is the door to the external world, because there is no way to represent everything within a museum or within a benchmark, which is what’s used for evaluation. And when what went off the rails in AI research is these claims of generality basically required people to say, I’m making a benchmark to test how general this is, but you can’t. And so The solution here is to get very specific about what we’re automating, why we’re automating it, and what kind of test would show us that it is sufficiently precise and accurate at that task that we could rely on it in that context. So stepping away from this like it’s general purpose thing, which is a fantasy and it’s untestable and will never be reliable to, yeah, like I am, I appreciate weather models that use statistical modeling to come up with likelihoods of where that storm is going to hit. That is a fantastic use of statistical modeling, also called machine learning. That’s great. I appreciate automatic transcription. It’s gotten much, much better. And it’s gotten better because language modeling has always been a part of automatic transcription as language models got better. You still want to use it carefully. There’s places where you don’t want to rely on it, but fine. that’s, that’s useful. Machine translation likewise has gotten better. The T in chat GPT stands for transformer. And that’s a particular mathematical widget basically inside the neural networks that was developed at Google to improve machine translation. Again, you don’t want to rely on machine translation in cases where it has life or death or other important stakes and you are not in a position to actually check the output or the person that you are foisting the translated material on is not in that position, but still can be useful technology. But none of this is about changing my opinion on whether or not we have something called artificial intelligence that matches what that word evokes or that that’s even a good thing to be trying to build.
Seb (36:30)
Hmm. I mean, there are a few narratives, should we say, that come up and I’d love to sort put them to you and get your response. So the first of them is basically, whatever you thought, people say this all the time, whatever you thought about AI three months ago, it’s already out of date. know, like scrap your knowledge, scrap your thinking. If you’re basing it off, yeah, three months ago, you’re out of date. need to kind of… Get with the times. Is this an argument you’ve heard and one you find convincing?
Emily M. Bender (37:06)
yeah, and about my own writing. Like one of the things people say, well, Stochastic Parrots was a fair characterization and then they’ll pick some time in the past. Sometimes it’s when the paper was actually published. Sometimes it was even last year. But now with, and then insert new hype technology, clearly it’s out of date. And it’s like, no, Stochastic Parrots is a description of how language models work. People are making the language models bigger. They are also bolting other things on. So what’s a likely next word is a terrible way to do arithmetic. Right. And early on, it was very easy to show how embarrassingly bad chat GPT was at that. I don’t think that’s true anymore. And I doubt that they just put in lots and lots and lots of math problems in the training data. The sensible thing to do would be to put a little text classifier on the front, catch as many math problems as you can and send them to an actual calculator, which is a good way to use computers to do math, right? As opposed to a large language model. So we don’t actually know what’s inside of these things. It’s very easy to make a flashy demo. It’s very easy to spin up like things people are excited or at least maybe I’m now three months out of date, but there was a period of time where people were very excited about the vibe coding tools that would let you put together an app. And I’m still seeing some stuff about that. And those apps like sure look good. They look like apps, right?
Seb (38:29)
Mm-hmm.
Emily M. Bender (38:30)
But were they tested? Like, do they actually do the thing that you want them to do? Probably not. So I think, you know, if any of this were actually any good, there wouldn’t be this urgency to get everybody to use it. The tech would sell itself.
Is Claude Code really a software engineer?
Seb (38:52)
Yeah, I can certainly see that. mean, for me, the one visible example, kind of counter example, which does feel like in my world at least sold itself and is selling itself very convincingly is kind of Claude code primarily and co-work to some extent, but code which has come along and depending on who you speak to is at least doing a very, very good job at a very good impression, should we say, of a convincing software engineer. And so that’s the one that I guess has come along. And yeah, in my world, as I say, I’ve seen like uptake of that shift dramatically. I’ve seen organizations that might previously have been sort of skeptical. Are we going to getting used out of this being like, no, we need to, you we need to change the way we operate in response to this particular technology. That doesn’t make the technology intelligent. It doesn’t make it sort of AI, but it does seem to be the one notable example in my world where it’s sort of, yeah, people seem to be getting a lot of value out of this.
Emily M. Bender (40:13)
Yeah, so I have many, many hesitations there. And the first is you said good impression of a software engineer. You’ve said that your background includes software engineering. Is a software engineer?
Seb (40:23)
Not, not, yeah, I worked in products. I worked very closely alongside software engineers. I’m not a, I’m not a software engineer by training.
Emily M. Bender (40:28)
Okay, all right. Is a software engineer only tasked with producing code?
Seb (40:36)
No, I would say not. I mean, it depends on what organization you go into truthfully. Like there are some where that kind of is the job description. They’re not typically the best. Those who are really good have a broad
Emily M. Bender (40:49)
Right?
Seb (40:49)
arena.
Emily M. Bender (40:50)
But my point is basically there’s a whole bunch of software engineering around conceptualizing the project, conceptualizing the algorithm, figuring out how it fits in, doing documentation, making something that is maintainable over time. And we started this conversation talking about work-slot. And I hear over and over again stories of people who get sent a change for code review, and it clearly just came out of one of these synthetic systems. And then they write back and say, I’m not going to review this code if the person who wrote it, wrote it in quotes, person who’s submitting it didn’t review it. Like you do your own job. And then you’ll sometimes get a second go around where that person then had one of the agents review it. And it’s like, if you don’t actually have the people in the organization really working with the code at sort of the most fine grained level, you are just setting yourself up for technical debt disasters down the Right. The, if the people didn’t write it, then the people don’t understand it. And you’re just going to have this mess of code that maybe functions now, but something changes. have to go in and debug. Like it’s going to get messy is one thing. A second thing is, people will say, okay, well, senior engineers can use this because they’re in a position to check the stuff that came out. And,
Seb (42:08)
Yeah. Yeah.
Emily M. Bender (42:09)
and I agree, right. The more expertise you have, the better positioned you are to check output. And programming is one of the cases where. you can fairly easily create systematic tests in a way that we don’t have in many other use cases. mean, does it compile? That is cheap and easy to do. And this, think, is also part of how these systems are getting better because they can use that as a training signal. Does this compile? It’s something that can be run automatically. But where do you suppose senior experienced software engineers come from?
Seb (42:42)
Yeah, yeah, exactly.
Emily M. Bender (42:44)
So we are to the extent that we need software engineering as a job, and I think we do. Software is important across many, sectors. If we are cutting off the pathways to becoming software engineers, we are setting ourselves up for a workforce problem down the line. And then on top of that, this stuff is expensive.
Seb (43:05)
Yes. And I think that stuff is actually, and it’s a very important point. I see a, there’s this narrative at the moment of like, basically lay people off and essentially spend as much, often much more on, you know, API budgets, basically to send data
Emily M. Bender (43:25)
Mm-hmm.
Seb (43:26)
to these large language models. And so, yeah, there’s now this kind of strange illusion of, we’re, cutting costs, you know, but people are basically just reallocating. I used to pay some person to do this and now I’m spending in some cases, as much, not way more on paying open AI or Anthropic for the right to use that technology.
Emily M. Bender (43:52)
And what happens when OpenA Anthropic raise their prices? Because they can’t burn cash forever.
Seb (43:55)
Yeah. Yeah.
Emily M. Bender (43:58)
Surprised at how long they’ve managed to burn cash, but like that has to end at some point. So, you know, not saying that there’s no place for automation in software development. If you think about things like version control, that’s a really important kind of automation. Think about things like compilers, Like running your unit tests, all of that. There’s lots and lots of really important automation. but I do think that it’s worth looking at the step between my ID has, code completion that reminds me of how I spelled the different variables to, I’m starting to type something that’s pretty boilerplate and I’m getting the suggested completion and I’m in a position to check that that’s what I was about to type to I’m writing in the documentation and it’s going to generate the code for me. Like the further away you get from actually touching each piece of it. I think the bigger of a maintenance nightmare you’re creating.
Seb (44:53)
Yes, yeah, definitely. to your point, mean, well, in my view, people seem to massively, they look at software engineering and they extrapolate from there into this narrative of AI takes all jobs, even in software engineering, where there are large parts of the job that are quite codified in a way that is just not true, I would say, of most people’s jobs. It doesn’t have this sort set distinct language that only a certain subset of people know. it’s even under those conditions, software engineering jobs are not disappearing anytime soon, certain types, more junior jobs sometimes, and you pointed to the problems there. But in a lot of cases, it just shifts. It does just abstract the same people to kind of the next level up. You know, they are maybe spending less time. literally tapping the keys to write a particular code, but they’re spending more time on thinking about system design, more time reviewing sort of higher level decisions, whether or not that’s a good thing, you know, there are pros and cons.
Emily M. Bender (46:08)
Yeah. Yeah. And I do think that we’re going to see more and more people getting hired to fix this kind of code. That this is, know, to the extent that the code was necessary and then it stopped working because it was built in a way where nobody really understood it. You’re going to have, I think, a group of people who are specialists in fixing that. Just like there was a group of people in the late nineties who had come up as cobalt programmers who were in massive demand ahead of Y2K.
Seb (46:32)
Yeah.
Emily M. Bender (46:34)
Yeah.
Seb (46:35)
Yeah, I mean, I know a few engineers who really sort of coded themselves into a job by knowing some like weird language that nobody really
Emily M. Bender (46:44)
Yeah.
Seb (46:45)
knows anymore, but some bank somewhere is highly dependent on it. And yeah, fair to say they have a
Emily M. Bender (46:51)
Yeah, yeah.
Seb (46:53)
nice life.
Emily M. Bender (46:54)
And this actually, I think is a really nice case in point about how when we put automation into critical systems, then we end up needing to put a lot of work into maintaining that automation that can get harder and harder over time. so it, when you said like, it’s going to take all the jobs generalizing from software engineering. Part of the narrative is coming out of these AI labs, clearly the smartest people are the AI engineers. And so if you could automate that. then you can automate anything else because they put themselves at the top of the scale, which is wrong in many ways. But also I think that in places where we don’t yet have very much automation, it’s worth thinking very carefully about why we would want to automate and thinking long-term, right? If we build our systems around this automation, what does that mean for resilience? What does that mean for institutional memory, all these things where when you have people involved, and people are tricky, People have annoying needs, but we also are social creatures and we have learned how to sort of build systems where people can fit into roles and keep working together. So even as the people shift out, it’s okay. And anytime we are looking at something that’s going to damage one of those systems, we need to tread, I think, very carefully.
Seb (48:15)
Yes. And as you said that we are social beings. I think part of the very real problem at the top here is in some of these companies, some of these are not really social beings. think
Emily M. Bender (48:30)
Antisocial.
Seb (48:32)
it’s a very real kind of Silicon Valley dream.
Emily M. Bender (48:35)
Mm.
Control, scaling, and world models
Seb (48:36)
There’s a topic about control here. It’s basically humans inconvenience me because they have their own opinions and They didn’t want to come to the office post COVID and all of these kinds of things that makes tech leaders sort of resentful. And the dream becomes almost like I get to sit in my room and twiddle the dials and it’s total control. Everything is taken care of. Nobody disagrees with me. there’s none of the inconvenience of dealing with real people and their real views of the world. yeah, I think that’s a very real problem in the picture here of some of the people shaping the technology here. So we’ve touched on this idea of, okay, whatever you thought about AI three months ago is out of date, I guess what I’m hearing is, well, it’s still. fundamentally, architecturally, it’s still kind of the same technology. It might be a slightly better version of that same technology. It might have been trained a bit better. It might have new innovations that make it feel more intelligent, quote unquote. But fundamentally, it’s still the same thing under the hood. There seems to be this belief Some people call it the scaling narrative. We just make these models better by basically just feeding them more and more and more data. And somehow in this process, they spontaneously shift into some intelligent being, a step-three
Emily M. Bender (50:15)
step three profit, right?
Seb (50:18)
prophet. So the scaling narrative would seem to be this dominant thing, just get bigger and bigger and bigger, and then somehow it turns into something else. People have become understandably very skeptical of that idea and sort of shifted away from it. The latest buzzword that seems to be going around more and more now is this idea of world models. This idea that like, okay, well, fine. If we just have these systems that are just predicting the next word, then of course they’re not going to be intelligent. What we need is some higher level model that kind of maps out, I don’t know. flimpses into the matrix and sees how everything is connected and sort of maps things out in that way. So yeah, I’m curious how much world models have crossed your desk as a topic. And if so, any thoughts on that?
Emily M. Bender (51:11)
Yeah. So first on the scaling thing, what the scaling narrative gets us is normalization of mass surveillance and ever bigger data centers, which with ever bigger environmental impacts. Right? So when you say it’s the same thing under the hood, yeah, it’s the same ideology of we get to take all the data and we get to, you know, just grab all the water and, you know, put huge impacts on the electrical grid because scale. And so, so just wanted to say that in terms of world models, People like to argue that the large language models, so things that are trained on massive amounts of text, because the text is about the world, therefore the models of which words go next to which other words is a kind of a world model. So I haven’t heard people saying we have to add world models, and I’m worried that it’s like some kind of abstraction over those parameters. I think if you were building, again, specific focused automation, In many cases, you would build an ontology, which is a world model. What are the things that exist in the domain that we’re working in here? How do they relate to each other? And then you can ground a bunch of stuff. So if you think about the voice assistants, so Google Home, Siri, Alexa, least the earlier iterations of those did have ontologies underneath. And so you could ask it for a specific kind of information and it would go into the ontology and then use that to map out. Yeah, I think world models built sort of intentionally are important. I worry that anytime you slap an LLM based interface over one of these things, you are basically obfuscating the actual affordances of the technology because it’s going to sound much more flexible than it actually is. And then people can’t use it well.
China, the arms race, and the genie in the bottle
Seb (53:02)
Yeah, interesting. Another narrative this one plays out in the news. I don’t know if it necessarily plays out in the sense that it comes across people’s desk, but yeah, China seems to be a big part of the kind of geopolitical narrative as we can call it, which is basically arms race. know, if you believe that AI exists and if you believe it becomes this sort of all seeing all dancing AGI, if we don’t build it, China will. And if China does then, yeah. and end of world scenario.
Emily M. Bender (53:37)
Yeah, I mean, it seems to have been a narrative that was very effective with US policymakers, unfortunately, to basically the tech companies go to US policymakers and say, don’t you dare regulate us, because if you regulate us, then we aren’t going to be able to outpace China, and then we get the bad Chinese version of this. It’s all based on that same fantasy, right, that these things are going to combustion to consciousness and be aligned with the values, hopefully, of whoever built them. and American values are supposed to be better than Chinese values. I’d like to point out that China’s carbon emissions are actually diminishing at this point, just as one value that we might think about trying to adopt here.
Seb (54:17)
Yeah, yeah, it’s interesting. mean, you know, they are very much portrayed as the bad guys, particularly by the US. In fact, my most recent guest on the podcast or the last but one was a lady by the name of Selena Xu. She’s a researcher of both AI and China. yeah, I mean, China has a bit of a different view of the kind of the AI future here and you can… people form their own view about whose they think is better, more realistic, whose values
Emily M. Bender (54:50)
Yeah. And, you know,
Seb (54:53)
they prefer.
Emily M. Bender (54:53)
the social scoring technology that we were hearing about from China a few years ago, that sounds like something I wouldn’t want to participate in. But the fact that people have built it in China can be problematic for people there. It’s not necessarily going to happen elsewhere. Like these things, despite being called agents, don’t have that kind of agency. Right. It’s all about what decisions do we make as people individually, collectively about what we’re going to automate and what we’re not going to automate and what we are going to regulate in the space of automation.
Seb (55:21)
Yeah. So, you know, some people listening might be great AI advocates and others great skeptics. Either way, there may be people out listening who think that essentially the genie is just out of the bottle now and you can’t put it back in. So what do you make of that particular argument?
Emily M. Bender (55:45)
I don’t buy it. So it’s true that the transformer architecture is not knowledge that we’re going to lose anytime soon. But the very large models that are driving so much of this, because you’ve got people directly accessing them, because you’ve got companies building skins on them and sort of providing other services and so on, are hugely expensive and not naturally occurring artifacts. And they don’t have to continue to exist. bubble might burst of its own, right? That might just become under its own weight, right? Or we might set up regulation. mean, there’s folks who’ve said, yeah, we wouldn’t be able to do this if we couldn’t just appropriate everyone’s data. Well, if we stood up and said, you know, you can’t appropriate everyone’s data, you have to destroy models that are built that way, they’re gone, right?
Pride in your own expertise
Seb (56:33)
Hmm. And I guess what’s your, what’s your view of the alternative future here? You know, is it, it that, the, technology actually, yeah, would you rather it just didn’t exist that it disappeared or do you think it should exist, with clearer parameters around what it can be used for? what’s, yeah, what’s your view of, of.
Emily M. Bender (56:56)
So I see no beneficial use cases for synthetic text. It is endangering information ecosystem. Basically, the killer app is fraud. So deep fakes, making it hard to find truthful accounts in media because all of sudden we have all these apparent news sites that are just posting synthetic text and on and on on like that. So yeah, I don’t think there’s a net benefit to the world for synthetic text. Language technology, sure. I run a professional master’s program in computational linguistics. The world that I want to move to is one where we have better regulation, where we don’t allow the kinds of accumulation of wealth that are really part of what’s driving this big bubble, and where technology is built in much more specific use cases and is much more under the control of the people who are being impacted by it. a message I’d like to leave your listeners with is have pride in your own expertise. Because the narrative coming from the tech companies that are selling so-called artificial intelligence is we can get a machine to do what you do. And that is always based on a misunderstanding of what it is that people are doing when we do our jobs.
Seb (58:05)
Yeah, I think that’s an important message that many people need to right now. So I appreciate you coming on and sharing your perspective. I really enjoyed this conversation.
Emily M. Bender (58:16)
Thank you and it’s great to talk with you and to your audience.





