Interactive Harappan script explorer
Indus Valley Script
Browse, search and analyse the undeciphered Harappan script of the Bronze-Age Indus Valley Civilisation — 715 catalogued signs and 5,445 inscriptions from 81 archaeological sites, complete with glyph images, find-spot maps and statistical tools.
Most frequent signs (this corpus)
Find-spots inscriptions per site
Largest sites (by inscriptions)
Text length (signs per inscription)
Chronology broad, approximate
Reading direction
Material
Object type
Completeness
| ID | Text | Site | Text code | Signs | Dir. | Period |
|---|
Rank–frequency (Zipf) log–log
Each point is a sign: its frequency rank (x) vs corpus frequency (y), both on log scales. A near-straight line is the Zipf pattern typical of natural scripts.
Positional profile reading order
How often each sign appears initial (first read), medial, terminal (last read), or solo (single-sign text). Reading order derived from the direction field (R/L assumed when unknown).
| Sign | # | Freq | % corpus | ICIT total | Initial | Medial | Terminal | Position |
|---|
Neighbours of a sign reading order
Top sign pairs bigrams
Top sign triples trigrams
Inscriptions per site top 25
| Site | Inscriptions | Distinct signs | Avg length | Top signs |
|---|
Distribution by field
Cross-tabulation
Sign co-occurrence matrix
Cell colour = how often the row sign and column sign occur together (darker = more). Hover a cell for counts. Diagonal shows each sign's own frequency.
Predictability of the next sign bits of uncertainty
Lower bars mean more predictable. If knowing the previous sign lowers uncertainty (conditional < unigram), the order carries information — a hallmark of language, and unlike a uniform-random script.
Per-symbol (block) entropy n = 1, 2, 3
Per-symbol entropy of n-grams divided by n. A moderate decline (not flat, not collapsing) is typical of structured language rather than random or rigidly fixed sequences.
Collocations pairs that co-occur more than chance
Adjacent sign pairs (reading order) ranked by association strength — candidate compound units / “words”, not just the most frequent pairs.
| Pair | Count | PMI | G² |
|---|
Recurring sign blocks candidate words / phrases
Contiguous sign sequences that repeat across many different inscriptions.
Positional distribution relative position in text
Each row is a sign; columns run from the start (left) to the end (right) of the inscription in reading order. Darker = a larger share of that sign's occurrences falls in that position band.
Distributional similarity “signs used like this one”
Signs whose neighbour contexts (what precedes/follows them) most resemble the chosen sign's — candidate functional classes.
Clusters of frequent signs grouped by shared context
Distinctive signs keyness (log-odds z-score)
Signs over- or under-represented in the selected group versus the rest of the corpus (weighted log-odds; higher |z| = more distinctive).
Over-represented here
Under-represented here
Identical inscriptions same sign sequence
Near-identical inscriptions differ by one sign
Sign transition network
Signs placed on a circle (by frequency); arcs link signs that frequently follow one another — thicker = stronger association.
Find-spots map approximate coordinates
Marker area is proportional to the selected measure. Click a site to see its inscriptions. Coordinates are approximate; a few minor find-spots without reliable coordinates are omitted.
About
The Indus script (Harappan script) is a corpus of symbols produced by the Indus Valley Civilisation. Despite many attempts it has not been deciphered; there is no known bilingual inscription, and the script shows little change over time.
This is an interactive browser for the sign inventory and the catalogue of inscribed texts, with tools to search, compare and analyse the script.
How to read a text
Each text's text code lists its signs left-to-right, separated by
-. A / starts a new line. Edges marked + are complete;
edges marked [ or ] are broken/damaged. Sign 999 denotes
an unknown/illegible sign and has no glyph.
References
- Interactive Corpus of Indus Text (ICIT) (ISBN 978-1-84217-994-9) by Bryan Kenneth Wells. Finished in 2006.
- ICIT Database of Indus Writing developed by Bryan K. Wells and Andreas Fuls as a co-operative effort.
Text corpus & SVG sign vectors from the original indus-valley-script project; canonical sign glyph images and sign frequencies (total & per-site) are credited under References above.