Case studies

A Century of Writing, Searchable by Who Said What

A think tank held the written record of a tradition going back to the 1850s: more than two thousand PDFs, most of them photographs of printed pages, on a site that was slow to load and impossible to search. The collection existed. The knowledge in it did not.

1,000+primary works catalogued
2,000+source PDFs, mostly page scans
506writers identified and classified
5languages searchable

First, Making the Pages Readable at All

A scan is a picture. A search engine cannot read a picture, and neither can software that wants to know who is mentioned on page forty. So the first pass turned the pictures into text, across five languages and several scripts, including printed material old enough to have unusual typesetting.

That gave us words. It did not give us knowledge, and most digitisation projects stop at that point.

BeforeHow it worked

Two thousand PDFs, most of them photographs of printed pages, on a site that was slow to load and impossible to search.

AfterHow it works now

Ask who wrote about whom, on what subject and in what words. Every relation carries the passage that proves it.

The Ontology Is the Part That Mattered

Layer 1Works

What kind of document each item is: a speech, a pamphlet, a book, an article in a periodical. Which run or series it belongs to, and where in that run.

Layer 2People

Every writer, resolved to one identity across scripts, spellings and initials, so the same person is not counted as several, however a byline was printed.

Layer 3Relations

Who wrote what, who is the subject of what, who is mentioned inside what, and who was arguing with whom.

Layer 4Positions

Where each writer sits in the tradition, on several independent axes, judged from what the collection actually shows about them.

Most archive projects build layer 1 and call it a catalogue. Layers 3 and 4 are what turn a catalogue into something you can ask questions of.

One Person Is the Hardest Part

A hundred years of printing produces a dozen spellings of the same name. Initials expand and contract, transliteration conventions change, a periodical drops a middle name, and the same writer appears in Devanagari in one work and Latin script in another. Until those collapse into one identity, every count in the archive is wrong.

  • The name as printed on a 1954 pamphlet
  • The initials-only form used by a periodical
  • The same name in a different script
  • A transliteration with different vowels
  • A byline with the middle name dropped
Resolved toOne writer

One identity, carrying every form it has ever been printed under, so a search on any of them finds all of the work.

Where the software cannot make the match with confidence, it does not. The name is kept exactly as printed and marked unresolved, which leaves a curator a real queue to work through. Forcing an uncertain match would quietly merge two people, and nobody would ever find the error.

Every Relation Carries Its Own Evidence

A statement like "this author engaged with that thinker" is worthless on its own, because you cannot check it. So no relation in this archive exists without the words that prove it. Each one records what kind of link it is, how confident the reading was, and the passage it came from.

A 1990 booklet on India's economy cites a classical economist Evidence held with the link: the sentence warning that India must act to avoid “the disaster which Malthus had forecast for a nation multiplying itself unchecked.”
The same booklet invokes the founder of a free-enterprise forum Evidence held with the link: the epigraph printed inside the front matter, attributed to him by name and dates.

The rule that makes this trustworthy is mechanical. Every quotation attached to a relation must appear verbatim in the text of the work it is drawn from. The pipeline checks each one as a substring, and any relation whose quotation fails that check is discarded rather than kept with a warning. A fabricated quotation cannot survive the step that stores it.

Positions, on Axes That Stay Separate

The interesting question about a tradition is not who is in it, but how. So each writer is placed on several independent axes rather than given one label: how central they are to the collection, which strand of thinking they belong to, and what they actually did for a living. Keeping the axes separate matters, because a person can be peripheral to the archive and central to the century, and a single tag would force us to pick.

Each placement is recorded with the confidence behind it, and anything uncertain is flagged for a curator instead of being quietly published as fact.

Held per writerThree separate readings
  • How central to this collection
  • Which strand of thinking
  • What they did: economist, editor, parliamentarian, industrialist
Held alongsideThe audit trail
  • A confidence level per axis
  • The reasoning, in a sentence or two
  • A review flag when any axis is uncertain

Roughly three quarters of the writers came out of the first pass flagged for human review. That is the system working. A classifier that returned confident answers for all of them would be lying about a corpus this uneven.

What the Method Turned Up

Before running the classification across the whole collection, we tested it against a set of answers written by hand. It disagreed on several writers. When we examined the disagreements, the software was right and the hand-written answers were wrong: it had counted how often each writer was actually written about in the collection, and we had gone on reputation.

We corrected our own answer key and let the run proceed. That is the argument for building the ontology from the corpus rather than from what everyone already believes.

Two Tiers, and the Archive Says Which Is Which

Old scans do not all read cleanly, and a confident quotation drawn from a badly read page is more damaging than no quotation at all. So the archive is explicitly split, and says which tier it is answering from.

Tier ATrusted text

Clean material with stable paragraph addresses. Searchable, quotable, and linkable down to the paragraph. Software may quote it directly.

Tier BScanned works

Full catalogue record, a summary and its main arguments, and the original document. Software may describe it and must link out for the underlying claim rather than quote it.

The tier travels with every record. Anything reading the archive can tell what it is allowed to say about a given work. The route from Tier B up to Tier A is already in the schema, so if a scan is cleaned up later, that is a change to the data rather than a rebuild.

What the Researcher Gets

A fast library, searchable in five languages. A name resolves to a person rather than a spelling. Every answer points back to the page it came from. And a question like who argued with whom about free enterprise in the 1960s has an answer you can follow down to the paragraph.

Software can read the same structure. A researcher's AI assistant works from the archive directly, under the same tier rules a person gets, and nobody has to hand over a copy of the collection.

Tell us your hardest problem. We will solve it.