The Brief
Society & Ethics 5 min read

A False Report, a Ship, and a Near-Miss: AI Hallucination in Wartime

NAVION

Share

An AI-generated intelligence report falsely identified a Chinese vessel as transporting components for a nuclear arms program. The US military came close to intercepting and boarding that ship, with air support ready. The error was caught before action was taken, but only barely. One source described the episode as having “almost started a war.”

This is not a hypothetical scenario about future AI risks. It is a reported near-miss that illustrates, with unusual clarity, what happens when AI hallucination collides with high-stakes decision-making.

How a Chatbot Rewrote a Ship’s Cargo

According to CNN’s reporting, a US Special Operations Command analyst used a chatbot to process intelligence about the Chinese vessel’s manifest. The tool combined open-source intelligence with classified signals intelligence held in government systems, then packaged the result into a formal report. That report concluded the ship was carrying nuclear arms program components through the Middle East.

The conclusion was entirely false.

The chatbot had, in the language now familiar to anyone following AI development, hallucinated. It generated a confident-sounding output that did not reflect reality, because the underlying data did not provide sufficient context for an accurate answer. The military was preparing a physical interception before officials identified the error in the AI’s output.

This is what most coverage of AI hallucination misses: the problem is not that AI systems occasionally produce wrong answers in low-stakes settings. The problem is that the outputs are formatted, structured, and delivered in ways that look authoritative. An intelligence report generated with AI assistance does not arrive with a disclaimer. It arrives looking like an intelligence report.

A Military That Is Accelerating, Not Pausing

The episode sits inside a broader pattern of rapid AI adoption across the US defense apparatus. In January, the Department of Defense launched what it described as an “AI acceleration strategy,” with the stated goal of making data available across federated systems for what the department called “AI exploitation.” Defense Secretary Pete Hegseth framed the initiative around data quality: “AI is only as good as the data that it receives.”

The department has moved quickly on the tools side as well. Last December, it announced the use of Google’s Gemini for Government as the foundation for a bespoke platform called GenAI.mil. The following month, Grok for Government was added as an option on the same platform. Anthropic offers a customized version of Claude for US intelligence work. A Pentagon representative told Congress in June that 1.5 million active Department of Defense personnel have used the military’s generative AI tools, and that generative AI is already being used to help produce congressionally mandated reports.

The scale of deployment is significant. So is the speed. The near-miss with the Chinese ship did not occur in a vacuum: it occurred inside an institution that has made AI integration a strategic priority.

There is a tension here that deserves attention. A 2023 State Department declaration on responsible military use of AI called for “careful consideration of risks and benefits” and insisted that accountable use of AI systems must always involve “a human in the loop, a responsible human chain of command and control.” The reported episode suggests that a human was technically in the loop, in the sense that an analyst submitted the report. But the analyst appears to have trusted the chatbot’s output without sufficient verification. The human in the loop is only as effective as their ability to interrogate what the AI has produced.

What This Reveals About AI in High-Stakes Environments

The hallucination problem is not new, and it is not confined to military contexts. Since “hallucinating” became the Cambridge Dictionary’s word of the year in 2023, documented cases have emerged across medicine, law, journalism, academic research, and corporate settings. Researchers have suggested that it may be impossible to prevent large language models from hallucinating entirely, regardless of how prompts are constructed.

What the military near-miss adds to this picture is a sense of consequence at a different order of magnitude. A hallucinating AI in a customer service context produces a wrong answer that can be corrected. A hallucinating AI in an intelligence context can produce a wrong answer that triggers a chain of events involving armed forces, international relations, and the possibility of armed conflict.

The deeper issue is structural. AI tools are being integrated into workflows faster than the verification frameworks needed to catch their errors are being built. The Department of Defense’s acceleration strategy is explicit about speed and scale. The safeguards, by contrast, are described in general terms: principled use, human oversight, careful consideration. These are not the same as operational protocols that would have caught the error in the Chinese ship report before it reached the stage of military preparation.

AI augments human analysis. It can process volumes of data that no individual analyst could handle alone. That is a genuine capability. But augmentation requires that the humans working alongside these tools have both the time and the institutional support to verify outputs critically, especially when those outputs will inform decisions with irreversible consequences.

In Short

An AI chatbot fused open-source and classified intelligence and produced a false report that nearly triggered a US military interception of a Chinese ship. The error was caught, but the episode reveals a structural gap: the US military is deploying AI at scale and speed, while the verification frameworks needed to catch hallucinations in high-stakes contexts are still catching up. AI hallucination is a known, documented, and possibly irreducible problem. In low-stakes settings, it is an inconvenience. In military intelligence, it is a different category of risk entirely.

Based on reporting from Ars Technica.

Written by

NAVION