The Brief
Science & Discovery 5 min read

AI Can't Decipher Lost Languages Alone. Here's Why That Matters.

NAVION

Share

For over a century, two ancient writing systems have resisted every tool linguists could bring to bear. Linear A, the script of the Bronze Age Minoan civilization on the Greek island of Crete, remains entirely undeciphered. Etruscan, spoken in Italy before the rise of Rome, is only partially understood. Both have now attracted serious attention from artificial intelligence. The results so far reveal something important: not about what AI can do, but about the precise boundary between pattern recognition and meaning.

The Hypothesis Comes First, the Machine Comes Second

A recent case illustrates the dynamic clearly. In June 2026, a self-taught AI engineer and amateur linguist proposed a potential breakthrough on Linear A. The starting point was a single human guess: that one unknown word in a prayer inscription might derive from a Semitic root meaning “to dwell” or “to inhabit.” From that hypothesis, AI-built programming scripts were used to test the proposed sound pattern against a collected corpus of Linear A characters. The result was an assignment of values to 40 signs and a compiled lexicon of 408 words, suggesting, in the researcher’s interpretation, that Linear A belongs to the Semitic language family, which includes Hebrew and Aramaic.

This is what most coverage of AI and ancient languages gets wrong. The AI did not generate the idea. It tested one. The distinction matters enormously. What the technology contributed was speed: cross-checking a human hypothesis against thousands of characters in a fraction of the time a manual review would require. That division of labor, human intuition directing machine verification, is the pattern worth understanding here, not the specific claim, which remains under review.

Where the Ceiling Is, and Why It’s Real

AI brings genuine strengths to this kind of problem. Large-scale pattern testing is one of them. A model can scan an entire archive for repeated sequences that a human eye would miss, and it can restore damaged or fragmentary inscriptions by predicting likely missing characters. There is also a technique called cross-lingual transfer, where a model trained on a known language can sometimes infer structural patterns in a closely related unknown one. The analogy is intuitive: knowing Spanish makes Portuguese partially readable. Researchers have applied this approach successfully to Ugaritic, an extinct Semitic language spoken during the late Bronze Age, roughly between 1300 and 1190 BC, in the ancient coastal city of Ugarit in what is now Syria. Because Ugaritic’s language family was already known, the anchor existed.

Linear A and Etruscan are harder precisely because that anchor is missing or incomplete. Statistical pattern matching cannot manufacture meaning from nothing. It needs a reference point, a known language family or a bilingual text, to distinguish meaningful patterns from coincidence. The entire surviving corpus of Linear A is approximately 7,500 characters, short enough to fit on a single screen. With that little data, almost any hypothesis can find scattered matches to support it. This is not a limitation of current AI models. It is a structural problem that no increase in computing power resolves.

A model could theoretically become fluent enough in Linear A to recognize repeating patterns and structural units, even approximating a limited form of interaction on its own statistical terms. But fluency and meaning are not the same thing. The model can learn which signs follow which, and which words cluster together, without ever knowing what any of them actually refer to. Handing a human a translation requires something the model does not have access to.

Why This Tells Us Something Broader About AI’s Role

The decipherment problem is a useful lens for thinking about AI’s actual position in knowledge work. The technology accelerates. It compresses years of manual cross-referencing into minutes. It allows more people to attempt problems that institutional resources previously kept out of reach. These are real contributions.

But acceleration is not the same as resolution. The two ingredients that decipherment has always required remain unchanged: a genuine comparative anchor, and rigorous human review capable of distinguishing a real breakthrough from an appealing coincidence. AI does not supply either of those. It operates between them, doing the volume work that humans cannot do at scale, while humans retain the judgment about what the patterns actually mean.

Verifying any AI-assisted claim about an undeciphered language is also structurally difficult in a way that has no easy fix. Normally, a proposed translation would be checked against native speakers, other texts, or expert consensus built over decades. None of that exists for Linear A or Etruscan. The only available check is independent expert scrutiny and peer review. That is why “AI found a pattern” and “AI found the correct meaning” are very different claims, and why the gap between them is easy to blur in coverage that focuses on the announcement rather than the epistemology.

In Short

AI is a genuine accelerant for the study of undeciphered languages. It can test hypotheses at scale, spot patterns across large corpora, and make these problems accessible to more researchers. What it cannot do is supply the anchor that decipherment requires: a known linguistic relative, a bilingual text, or a verified comparative framework. Linear A and Etruscan remain unsolved not because the tools have been insufficient, but because the foundational reference point does not yet exist. Until it does, AI’s role in this field is that of a very fast research assistant working on a problem that still needs a human to frame the right question.

Based on reporting from Ars Technica.

Written by

NAVION