Indus Script: A Falsification-Oriented Analysis

A machine built to distrust its own best result. On two independently digitized versions of the Indus inscriptions, a promising statistical signal reverses sign once duplicate inscriptions are kept out of the test folds.

One result, three ways of scoring it

Predictive gain, in bits, from knowing the two preceding signs instead of one. Positive means the extra sign helps predict the next one.

Ordinary cross-validation

Corpus 1
Corpus 2

Duplicates grouped, so a sequence never appears in both training and its own test

Corpus 1
Corpus 2

Duplicates collapsed to unique sequence types

Corpus 1
Corpus 2
Values are from the paper's abstract. The two corpora are two independent digitizations of the same underlying corpus tradition, one of them the M77 concordance behind the most cited statistical studies of the script.

This project does not decipher the script, and it makes no claim to have read a single sign. Its question is narrower: which statistical claims about the script survive being attacked?

Why the signal flips

Between 23 and 33 percent of the sequences in both corpora are exact repeats of another sequence, which is what you would expect from seals stamped many times. That repetition is a real property of the ancient material, not a data error. But when repeats sit on both sides of the training and test split, a model can look good by recognizing a sequence it has already seen.

A control that grouped sequences with the same size distribution, but without duplicate identity, stayed strongly positive. So the reversal is not simply a smaller effective training set. It tracks duplicate identity specifically. The paper treats resolving this as its central open problem.

What else the tests found

The signal does not point to language
The ordinary-CV gain is not uniquely diagnostic of the linguistic generator tested, even after rebuilding that generator and a non-linguistic alternative to match the real corpus's vocabulary and length statistics. Neither reproduces the real magnitude.
Small samples are unreliable
On two calibration languages, the estimator rarely returned a positive result below roughly 3,000 sequences. That is evidence against the ordinary-CV result being generic small-sample behavior.
Not a serial-number system
There is evidence against inscriptions being deliberately constructed as globally unique identifiers, though not against more elaborate identifier schemes.
A bug the tests caught
One digitization had its sequences reversed. The check that exposed it used a previously reported pattern: the last sign in a line is more predictable than the first.

How would we know if someone did crack it?

A companion article turns the same attitude on decipherment claims. It proposes three questions, asked in order, and a claim has to clear each before the next is worth asking.

  1. Is there any language-like signal?Do the sign sequences show statistical structure that a natural language would? This is necessary and far from sufficient.
  2. Does that signal survive proper testing?Duplicate-aware evaluation, sample-size matching, and robustness to how signs are grouped. This is the level the toolkit builds, and where the sign flip above lives.
  3. Does a specific reading survive independent evidence?Its proposed readings must make risky predictions on evidence that did not shape them, as Ventris's Linear B readings did when the Pylos tripod tablet turned up in 1953.

The article's reading of the record is that none of the published Indus proposals has yet passed the third test, and that this is not the same as the script being unreadable. It ends with a sentence worth keeping: not yet, but here is exactly what would change my mind.

The gap: what an actual decipherment would need

Statistics can tell us what kind of structure is in the sign sequences. It cannot, by itself, tell us what the signs say. Three facts about the evidence explain why the gap is so wide.

Two next steps follow directly from the work so far. They are proposals, not results.

Read and reuse