You turn to us for voices you won't hear anywhere else.

Sign up for Democracy Now!'s Daily Digest to get our latest headlines and stories delivered to your inbox every day.

“Anatomy of an AI Kill Chain”: Militaries Rely on Mistake-Prone AI in Ukraine, Gaza & Iran

Listen
Media Options
Listen

Militaries across the world are in an “AI arms race,” says Heidy Khlaaf, a chief AI scientist at the AI Now Institute, which just released a joint report with Airwars titled “Anatomy of an AI Kill Chain.”

“What it is is a series of different types of AI algorithms that are chained together, each with their own pitfalls, each with their own flaws,” she says. “And these systems build off of each other and feed into one another to, unfortunately, end up resulting in civilian casualties.”

This can lead to faulty targeting because “humans are prone to just accepting recommendations of AI algorithms without corroborating that evidence,” adds Khlaaf.

Related Story

StoryJul 08, 2026NATO Meets in Turkey Amid Crackdown on Civil Society; Trump Praises Erdoğan & Considers F-35 Sales
Transcript
This is a rush transcript. Copy may not be in its final form.

AMY GOODMAN: This is Democracy Now!, democracynow.org, The War and Peace Report. I’m Amy Goodman in London, with Nermeen Shaikh in New York.

NERMEEN SHAIKH: We turn now to how artificial intelligence is transforming warfare, from targeting software to autonomous drones. Earlier this year, Defense Secretary Pete Hegseth spoke at Elon Musk’s company SpaceX.

DEFENSE SECRETARY PETE HEGSETH: In short, we will win this race by becoming an AI-first warfighting force across all domains, from the back offices of the Pentagon to the tactical edge on the frontlines.

NERMEEN SHAIKH: And this is CENTCOM Commander Admiral Brad Cooper talking about the U.S. military’s use of AI.

ADM. BRAD COOPER: Our warfighters are leveraging a variety of advanced AI tools. These systems help us sift through vast amounts of data in seconds, so our leaders can cut through the noise and make smarter decisions faster than the enemy can react. Humans will always make final decisions on what to shoot and what not to shoot, and when to shoot, but advanced AI tools can turn processes that used to take hours, and sometimes even days, into seconds.

AMY GOODMAN: We’re joined now by Heidy Khlaaf, the chief AI scientist at the AI Now Institute. It’s just released a joint report with Airwars titled “Anatomy of an AI Kill Chain.”

Heidy, thanks so much for being with us. We want you to explain this report. But also, to warn you, to a global lay audience, when you use terms like “generative AI” and “artificial general intelligence,” if you would explain your terms? But tell us about this “Anatomy of an AI Kill Chain.”

HEIDY KHLAAF: Yes, we actually wanted to break down and demystify a lot of these terms that we’re now seeing being used, like you mentioned, “generative AI,” “large language models,” and also, more generally, the use of AI generally in warfare and what that means.

And I think it’s important to remember, and what we’re trying to convey in this report really, is that AI has been used in the military and defense since the 1960s. And what we’re — when we’re talking about a new AI arms race, it means the inclusion of something like ChatGPT, which is a type of AI that’s known as a large language model, into now the decision-making processes for targeting. Right?

And the thing is, that we want to break down, is that it’s not just one AI killer robot, which is often kind of the concept that people have of AI. What it is is a series of different types of AI algorithms that are chained together, each with their own pitfalls, each with their own flaws. Each have very low reliability and accuracy. And these systems build off of each other and feed into one another to, unfortunately, end up resulting in civilian casualties, as we’ve often seen with their use.

So, we break down each step of the kill chain, how each AI algorithm leads to faulty decisions and ultimately can lead to decision-making in targeting civilian casualties because of how faulty these algorithms are. We don’t just focus on the new types of AI that we are seeing advertised, you know, as discussed, things like ChatGPT, large language model. We also talk about older types of AI that continue to be used and continue to be relied on, like vision systems.

NERMEEN SHAIKH: And if you could explain, Heidy, the difference between decision support systems and autonomous weapons systems? And on the question of large language models, you’ve warned that there have been operations from Russia and China that put out propaganda to try to skew the outputs of large language models. What does that mean?

HEIDY KHLAAF: So, really, the difference between decision support systems and autonomous weapons systems is that there’s a human in the loop when you’re selecting a target, essentially, to strike. An autonomous weapons system means that this is done automatically without a human in the loop.

And I consider this kind of, you know, separation between them to be pretty superficial in practice, because, ultimately, if you have a human who is just rubber-stamping the AI decisions, right? And we know that they do this due to kind of something called automation bias in our field, which is that we know that humans are prone to just accepting recommendations of AI algorithms without corroborating that evidence. To me, in practice, that means that you have a human who, under the time pressures of war and conflict, is just approving decisions by an AI. And so, typically, this separation is often made to say, “No, the AI is safe. It’s fine even if it makes mistakes, because you have a human correcting that.” But in practice, we see a very, very different photo, or a different picture, really, of how that occurs.

NERMEEN SHAIKH: If you could explain also, Heidy, the way in which AI systems were used both by Israel in Gaza, as well as in the U.S.-Israeli war on Iran, and, in particular, the attack on the school in Minab in the first days of the war?

HEIDY KHLAAF: Yeah, I mean, we’ve known that large language models have been used since, really, the IDF started using them in Gaza. A good example of that is an AP investigation which demonstrated that they were using it to translate intercepted communications. Right? And they were taking, essentially, Arabic and determining whether or not specific people should be put on a target list. And we found, essentially, that because, again, large language models are highly inaccurate and unreliable, that they mistranslate Arabic. In this case, they took something like a “payment” and thought it referred to a “payload,” and added people to a targeting list. In this specific example, we know that this was caught. But again, under the time pressure, there’s going to be oftentimes situations where human operators aren’t going to double-check that, especially that oftentimes it’s very difficult to understand why a specific AI made the decision that it did.

And now we know that the — you know, it’s been confirmed by the U.S. Department of War that they are using Claude, through Palantir’s Maven, to similarly make targeting decisions. And we have now, unfortunately, the Minab school tragedy, which resulted in 160, you know, civilian casualties, and it’s becoming then unclear whether or not AI is used, that we know that Claude is being used to sort of make these targeting recommendations, but, ultimately, it hasn’t really been clear: Was it the AI? Was it not? Was it deliberate? Was it not? And how do we really trace that when we don’t really — when we can’t really verify or validate whether something was an AI recommendation or not?

Furthermore, we have a very difficult time with large language models because of their scale and nature of operation to be able to understand why an AI made the decision that it did. Was it due to faulty data? Was it directed to do so? Was it due to hallucinations? Which, if anyone has used ChatGPT, knows that these, you know, chatbots often kind of fabricate information. And I think that’s where we are with that current situation. It’s still unclear to us what was the cause of that. And I would say, actually, that’s kind of the purpose of AI systems themselves: They are often used to launder accountability.

AMY GOODMAN: Heidy Khlaaf, Anthropic earlier this year announced it’d be ditching its core safety promise. You previously worked at companies like OpenAI, where you tried to help them create their first safety frameworks to understand how to evaluate their AI systems. And now OpenAI’s head of ethics left after less than a year after joining. That was Chloé Bakalar, one of several high-profile exits at OpenAI in recent weeks as concerns mounted. If you can talk about what these ethics frameworks are and why people are leaving, one after another, and why it is it’s only the company that determines this, not government regulators, how AI is used?

HEIDY KHLAAF: Well, I can’t necessarily speak on why those specific individuals left, but I can speak from my own experience as a safety engineer, because I, prior to joining OpenAI, actually worked in nuclear defense and aviation, ensuring that the systems implemented in those, you know, what we call safety-critical situations don’t cause any harm.

And ultimately, what these companies are are they’re reinventing safety. They’re coopting a lot of the terms that we’ve used traditionally in safety and security engineering to really make this about kind of different types of safety. Instead of being concerned, “Do our systems — are our systems reliable? Are they accurate? Do they perform as intended? Are they fit for purpose, so that no human is harmed in sort of — due to the decisions that they or recommendations that they make?” we instead have these, like, safety and ethics frameworks about, you know, whether or not these systems are autonomous, whether they have intention, whether they’re conscious. Right? And that’s a very, very different question. And often they can give the illusion that they are, and sort of, you know, personifying a lot of the language, you know, or the actions of these models themselves. I mean, we’ve seen recently, you know, the cybersecurity capabilities being conveyed as kind of like, “Oh, this is autonomous. This is something we have expected from these things. They’re doing this on their own will.” When these companies actually deployed them in a very insecure environment, they intentionally were having these models, you know, go on to sort of carry out cyberattacks as part of sort of testing or benchmarking.

And so, this is — this is the reason that I take great issue with that, is that they actually, again, try to assign accountability to the AI rather than themselves. Really, I view that AI only does what we tell it to do. It’s only capable of what we build it to do. And these AI companies are, again, moving away from accountability by trying to assign sort of autonomy or personify these AI themselves in doing that. And I think, you know, when you’re then deploying these kind of models in warfare, it’s very, very easy to then sort of say, “Well, it wasn’t us who made that decision; it was AI.” And that’s why I take great issue with the way they’re now talking about safety and security engineering.

NERMEEN SHAIKH: And indeed, as you say, I mean, now it’s very routine to hear about AI agents, so to speak, going rogue. And yesterday there was a report that Taiwan was hit by an especially abnormal AI-assisted cyberattack, suspected to be from China. And reportedly, this kind of end-to-end autonomous attack on a government target has never been seen before. And the most remarkable fact about the target, as the Financial Times wrote, is that the AI continuously ranked and reprioritized possible attack plans based on available evidence. So, if you could talk about whether you think, in all of these cases, or the ones that we know about at least, it is the makers of this AI who are responsible, and not the agents that are going rogue, and what it would mean if these agents, artificial intelligence agents, move beyond cyberattacks and into actually theaters of war?

HEIDY KHLAAF: Well, the thing is, we are seeing them being deployed in warfare already. And the thing about AI is that it is a very powerful technology. They are very capable things, but they’re only as capable as we make them, and they only really do the things that we make them do. And oftentimes we might not be precise in the way that we specify their tasks, right? Like, the cybersecurity attacks, they were assigned to go and carry out cybersecurity attacks. You know, they were not doing this on their own will. They were built that way. They were trained that way. And I think that’s very important to keep in mind.

And in warfare, it’s the same situation. They are being basically given a lot of data to understand how to target. They are being given profiles, often faulty ones, to determine who’s an adversary and who isn’t. And ultimately, the thing is, these are based on really loose parameters, right? These are not — the way that we build them isn’t always accurate, isn’t always reliable. And we know that AI, really, doesn’t really understand anything beyond what it’s trained on. So, as soon as it sees a situation, especially in warfare, that it does not recognize, we know this is the exact situation, that the AI will fail, make a recommendation that’s not true, or fabricate outputs, right?

And so, I think it’s really important to keep in mind, yes, they are capable. Yes, they can do things like cybersecurity attacks. And they do speed up operations. They definitely do. That’s kind of, you know, one of their core features. But then, if you have something that speeds up operations, but their kind of baseline accuracy and reliability is 30 to 60%, that’s not really different from indiscriminate targeting, right? It doesn’t also mean this is some sort of, you know, killer robot that has its own intention. Anything has been built by design. And if we’re reckless in the way that we build them, that doesn’t mean it’s the AI is at fault. It’s — I think, definitely, the accountability should always be with the system designers and the people choosing to deploy them, despite knowing that they have these really poor reliability and accuracy rates, especially in new situations, which is actually more likely to happen in warfare than it is with sort of cybersecurity attacks.

AMY GOODMAN: Heidy Khlaaf, we want to thank you for being with us, chief AI scientist at the AI Now Institute. We will link to your report that you released jointly with Airwars, titled “Anatomy of an AI Kill Chain.”

Coming up, Timnit Gebru, the founder of the Distributed Artificial Intelligence Research, fired from Google in 2020 for warning about biases built into AI models. Stay with us.

[break]

AMY GOODMAN: “Dark Matter” by The Corner Laughers.

The original content of this program is licensed under a Creative Commons Attribution-Noncommercial-No Derivative Works 3.0 United States License. Please attribute legal copies of this work to democracynow.org. Some of the work(s) that this program incorporates, however, may be separately licensed. For further information or additional permissions, contact us.

Next story from this daily show

“Deep Unlearning”: Timnit Gebru on AI Hype, Ethics & Algorithmic Racial Bias

Non-commercial news needs your support

We rely on contributions from our viewers and listeners to do our work.
Please do your part today.
Make a donation
Top