The internet collectively lost its mind when news broke that a neural network could translate 5,000-year-old cuneiform tablets into English. The jokes wrote themselves, finally, we could automate the translation of 4,000-year-old complaints about Ea-nasir’s substandard copper deliveries. (Yes, that’s a real historical figure. Yes, his customer service reviews were bad enough to survive five millennia. No, we still don’t know what happened to his reputation.)
But beneath the memes lies something genuinely significant: researchers from Ariel University trained a model that can translate digitized Akkadian cuneiform into English, opening the door to hundreds of thousands of previously unread texts. The original research, published in PNAS Nexus, represents a watershed moment for digital humanities, and a reality check on what AI translation can actually do.

A Script That Outlived Empires
Cuneiform isn’t a language, it’s a writing system that was adapted for at least 15 different languages, including Akkadian, Sumerian, Hittite, and Elamite. Think of it as the USB-C of ancient writing: one connector, many different protocols running through it.
The system emerged around 3400, 3300 BC and persisted until 75 AD. That’s an astonishing 3,400-year run, longer than the entire history of the Roman Empire, longer than the gap between the fall of Rome and today. The last securely dated cuneiform text was written when Nero was emperor, and the script had already been in use for longer than Christianity has existed.
The Akkadian language itself is one of the earliest known Semitic languages, related to modern Arabic and Hebrew. It was the lingua franca of Mesopotamia, used for everything from legal contracts and administrative records to epic literature and scientific treatises. The fact that we can read any of it is a miracle of 19th-century scholarship.
The Decipherment That Took Decades, And the AI That Learned It in Hours
Translating cuneiform isn’t like translating French. There’s no Rosetta Stone equivalent for Akkadian in the way most people imagine, actually, there kind of is. The Behistun Inscription, discovered in Iran and dating to King Darius I of Persia (550 BC), contains the same text in Old Persian, Elamite, and Akkadian cuneiform. Deciphering Old Persian first provided the linguistic key to unlock the others.
By 1857, confidence in cuneiform decipherment was high enough that the Royal Asiatic Society sent the same Assyrian inscription independently to four scholars as a test. They all produced comparable translations. But that 1857 milestone was just the beginning, scholars have spent the 170 years since refining readings and interpreting difficult passages.
Here’s where the backlog problem kicks in. Hundreds of thousands of cuneiform tablets have been excavated. Many remain unpublished, untranslated, or both. “There’s tons and tons of cuneiform tablets and it’s well understood”, as one commenter on r/programming noted. “AI should be able to hack this fine.”
They were right, with caveats.
How the Model Actually Works
Shai Gordin and colleagues from Ariel University trained two versions of their neural network. The first translates Akkadian that’s already been transliterated into the Latin alphabet, a scholarly representation of the signs and their likely phonetic readings. The second translates directly from Unicode cuneiform characters into English.
The transliteration-based model performs better, achieving a score of 37.47 on the Bilingual Evaluation Understudy 4 (BLEU4) metric. For context, the BLEU score measures how closely a machine translation matches human-created reference translations, ranging from 0 to 100. Even experienced human translators rarely hit 100, and for a language with the complexity of Akkadian, 37 is genuinely useful for a first pass.

But the model has real limitations. It struggles with longer sentences, where maintaining context becomes exponentially harder. It also suffers from a problem familiar to anyone who has wrestled with modern LLMs: hallucination. The model occasionally produces translations that are syntactically flawless but completely disconnected from the source text’s meaning.
There’s a particularly instructive example from the ZME Science coverage of the research:
Source: UD 21-KAM2 LUGAL ina E2-DINGIR E2-DINGIR la ur-rad
Human translation: “On the 21st day the king does not go down to the House of God.”
Machine translation: “On the 21st day the king goes down to the House of God.”
The AI missed a negation. A single dropped “la” transformed a statement about royal absence into one about royal presence. In an administrative text, that’s the difference between “no shipment arrived” and “shipment arrived.” Someone auditing ancient grain accounts would have very different conclusions depending on which translation they trusted.
Why the Critique of “AI Deciphered Cuneiform” Headlines Was Right
Plenty of pushback emerged in the Reddit discussion of this research, and much of it was warranted. One particularly sharp commenter pointed out that we’ve known how to translate Akkadian since 1857. “This AI didn’t figure out anything about the language we didn’t already know.”
That’s technically true. But it’s also somewhat beside the point.
The problem isn’t decipherment, it’s scale. A few hundred people worldwide can read cuneiform fluently. Hundreds of thousands of tablets sit in museum basements and university archives. The durability of clay as a writing medium means we have far more Akkadian texts than we have trained humans to read them. Even if every Assyriologist on Earth worked full-time on translation, it would take generations to process the existing corpus.
This is a scaling problem, not a comprehension problem. AI can’t replace human expertise, but it can triage. As another commenter put it: “I would assume that this will be used as a first pass so people that aren’t able to read cuneiform can sift through these texts to spot the more interesting ones that can then be properly translated by a human.”
That’s exactly how the technology has evolved since the original 2023 study.
The Pipeline Problem: Beyond Translation
Translation was only the beginning. Since the original research, AI’s role in Assyriology has expanded into what researchers call the “pipeline problem”: photograph a tablet, detect its wedges, identify the signs and language, reconstruct damaged passages, and produce a translation that clearly marks uncertainty.
No current system can complete that entire process reliably without substantial human supervision. But components are emerging independently:
- Sign recognition and damage reconstruction: New systems can identify signs on damaged tablets and reconstruct missing portions.
- Fragment assembly: In 2025, researchers Anmar Fadhil and Enrique Jiménez used AI to locate and assemble fragments of a previously unknown hymn to Babylon. The work survived across 20 manuscripts written between the seventh and second or first centuries BC. AI helped find the pieces, human archaeologists reconstructed, interpreted, and translated the text.
- Expanding beyond Akkadian: In 2024, researchers released SumTablets, a dataset pairing Unicode cuneiform signs with scholarly transliterations for 91,606 Sumerian tablets.
- Benchmarking progress: The 2025 EvaCun challenge tested language models on Akkadian and Sumerian tasks like identifying dictionary forms and predicting missing portions of damaged texts, less flashy than full translation but critical for workflow augmentation.

The most recent development came in July 2026 with TabletCraft, a bidirectional translation model trained on 116,000 Akkadian-English examples. Unlike the 2023 system, it can also attempt to translate English into Akkadian and render the result as cuneiform, a capability aimed at education and public engagement rather than scholarly research.
What This Means for AI Beyond Ancient Languages
For technologists, the cuneiform translation work offers a useful case study in AI application development. This isn’t a generic translation problem, it’s an extreme domain with unique constraints:
- Limited training data. Akkadian corpora are minuscule compared to modern language datasets.
- High stakes for accuracy. A missed negation changes historical interpretation.
- Hybrid human-AI workflows are non-negotiable. The technology’s value is in accelerating scholarship, not replacing it.
- The infrastructure problem matters. The AI hardware discussion, whether you’re debating the cost analysis of local AI inference or memory bandwidth as the limiting factor, filters down to every application, even in digital humanities.
The enterprise investment in local AI experimentation that’s reshaping software development is the same compute story playing out in archaeology, only with fewer GPUs and more clay.
The Ea-Nasir Problem, Solved at Last
Here’s the thing about that Ea-nasir meme: it’s not just a joke. The complaint tablets about his allegedly substandard copper are genuine administrative texts from ancient Ur. They’re exactly the kind of document that fills museum collections, mundane records of commerce, law, and daily life that never make it into history books but collectively form the backbone of our understanding of ancient economies.
AI translation won’t just unlock epic poetry and royal inscriptions. It will unlock the ancient equivalent of Yelp reviews, tax filings, and lease agreements. That’s not glamorous. But it’s how we build a fuller picture of what daily life was actually like 4,000 years ago, not just what kings wanted recorded.
The researchers themselves acknowledge the technology’s limits. “No current system can complete that entire process reliably without substantial human supervision”, they note. Translation accuracy, hallucination risk, and contextual understanding all need improvement.
But the trajectory is unmistakable. What began with four Victorian scholars comparing notes on a single inscription has become a neural network trained on tens of thousands of documents, working alongside humanists to reconstruct poems from scattered fragments, identify tablet joins across museum collections, and process texts that might otherwise remain unread for decades.
The security risks in autonomous AI agent systems that keep modern AI engineers up at night feel quaint compared to the challenge of making sure a 4,000-year-old ledger about barley shipments isn’t misread. At least modern prompt injection attacks can be patched. Ancient cuneiform damages and dialect variations require something closer to scholarly judgment.
The takeaway for AI practitioners is simple: the hardest translation problems aren’t the ones with the most training data, they’re the ones where errors compound into permanent misreadings of human history. Build the tools, but keep the humans in the loop. And maybe double-check the negations.

