A Century of Writing, Searchable by Who Said What
A think tank held the written record of a tradition going back to the 1850s: more than two thousand PDFs, most of them photographs of printed pages, on a site that was slow to load and impossible to search. The collection existed. The knowledge in it did not.
First, Making the Pages Readable at All
A scan is a picture. A search engine cannot read a picture, and neither can software that wants to know who is mentioned on page forty. So the first pass turned the pictures into text, across five languages and several scripts, including printed material old enough to have unusual typesetting.
That gave us words. It did not give us knowledge, and most digitisation projects stop at that point.
Two thousand PDFs, most of them photographs of printed pages, on a site that was slow to load and impossible to search.
Ask who wrote about whom, on what subject and in what words. Every relation carries the passage that proves it.
The Ontology Is the Part That Mattered
What kind of document each item is: a speech, a pamphlet, a book, an article in a periodical. Which run or series it belongs to, and where in that run.
Every writer, resolved to one identity across scripts, spellings and initials, so the same person is not counted as several, however a byline was printed.
Who wrote what, who is the subject of what, who is mentioned inside what, and who was arguing with whom.
Where each writer sits in the tradition, on several independent axes, judged from what the collection actually shows about them.
Most archive projects build layer 1 and call it a catalogue. Layers 3 and 4 are what turn a catalogue into something you can ask questions of.
One Person Is the Hardest Part
A hundred years of printing produces a dozen spellings of the same name. Initials expand and contract, transliteration conventions change, a periodical drops a middle name, and the same writer appears in Devanagari in one work and Latin script in another. Until those collapse into one identity, every count in the archive is wrong.
- The name as printed on a 1954 pamphlet
- The initials-only form used by a periodical
- The same name in a different script
- A transliteration with different vowels
- A byline with the middle name dropped
One identity, carrying every form it has ever been printed under, so a search on any of them finds all of the work.
Where the software cannot make the match with confidence, it does not. The name is kept exactly as printed and marked unresolved, which leaves a curator a real queue to work through. Forcing an uncertain match would quietly merge two people, and nobody would ever find the error.
Every Relation Carries Its Own Evidence
A statement like "this author engaged with that thinker" is worthless on its own, because you cannot check it. So no relation in this archive exists without the words that prove it. Each one records what kind of link it is, how confident the reading was, and the passage it came from.
The rule that makes this trustworthy is mechanical. Every quotation attached to a relation must appear verbatim in the text of the work it is drawn from. The pipeline checks each one as a substring, and any relation whose quotation fails that check is discarded rather than kept with a warning. A fabricated quotation cannot survive the step that stores it.
Positions, on Axes That Stay Separate
The interesting question about a tradition is not who is in it, but how. So each writer is placed on several independent axes rather than given one label: how central they are to the collection, which strand of thinking they belong to, and what they actually did for a living. Keeping the axes separate matters, because a person can be peripheral to the archive and central to the century, and a single tag would force us to pick.
Each placement is recorded with the confidence behind it, and anything uncertain is flagged for a curator instead of being quietly published as fact.
- How central to this collection
- Which strand of thinking
- What they did: economist, editor, parliamentarian, industrialist
- A confidence level per axis
- The reasoning, in a sentence or two
- A review flag when any axis is uncertain
Roughly three quarters of the writers came out of the first pass flagged for human review. That is the system working. A classifier that returned confident answers for all of them would be lying about a corpus this uneven.
What the Method Turned Up
Before running the classification across the whole collection, we tested it against a set of answers written by hand. It disagreed on several writers. When we examined the disagreements, the software was right and the hand-written answers were wrong: it had counted how often each writer was actually written about in the collection, and we had gone on reputation.
We corrected our own answer key and let the run proceed. That is the argument for building the ontology from the corpus rather than from what everyone already believes.
Two Tiers, and the Archive Says Which Is Which
Old scans do not all read cleanly, and a confident quotation drawn from a badly read page is more damaging than no quotation at all. So the archive is explicitly split, and says which tier it is answering from.
Clean material with stable paragraph addresses. Searchable, quotable, and linkable down to the paragraph. Software may quote it directly.
Full catalogue record, a summary and its main arguments, and the original document. Software may describe it and must link out for the underlying claim rather than quote it.
The tier travels with every record. Anything reading the archive can tell what it is allowed to say about a given work. The route from Tier B up to Tier A is already in the schema, so if a scan is cleaned up later, that is a change to the data rather than a rebuild.
What the Researcher Gets
A fast library, searchable in five languages. A name resolves to a person rather than a spelling. Every answer points back to the page it came from. And a question like who argued with whom about free enterprise in the 1960s has an answer you can follow down to the paragraph.
Software can read the same structure. A researcher's AI assistant works from the archive directly, under the same tier rules a person gets, and nobody has to hand over a copy of the collection.
