Collections That Answer Questions
An archive already holds the knowledge. What it lacks is a way for that knowledge to be asked a question, by a researcher or by anything else.
We have done work of this kind. See the think tank archive for how the mapping works in practice, and the open library at Falsafa.ai for the method in public.
The Path a Collection Takes
Scanning, photography of bound and fragile material, and recording. What exists only on paper or only in someone's memory becomes a file that can be worked with.
Pictures of pages become text, including old typesetting, many scripts and handwriting. Recordings become transcripts with speakers and timings.
The ontology: what each item is, who made it, what it is about, who is named inside it, and how items relate. Every relation carries the passage that proves it.
To people as a fast public platform, and to machines as an open interface. Both read the same structure.
Most digitisation stops after step 2. That produces a searchable pile, which is better than a cupboard and still not a knowledge system. Steps 3 and 4 are where a collection becomes something you can ask a real question of.
Oral History as a First-Class Source
Testimony is the most perishable holding any institution has, and the least well handled. We record it, transcribe it with speakers separated and timings kept, and translate where the interview is in a language the catalogue is not. Mixed-language speech is tagged honestly as mixed rather than forced into one language, because that tag decides how strictly a later quotation can be trusted.
The transcript is then treated exactly like a manuscript. People and places named in it become entities. Passages get a fixed address so they can be cited. What the speaker said about a subject becomes a relation, with the words attached.
Material the Machine Finds Hard
Older holdings break ordinary pipelines: unusual scripts, marginalia, damaged pages, mixed languages on one leaf, and layouts no modern reader expects. Two commitments make this survivable. The original stays untouched and byte-exact, so every later pass can be rerun from the source rather than from a processed copy. And the catalogue declares how confident it is in each reading rather than presenting a shaky transcription as fact.
An Open Interface, on Your Terms
Researchers now arrive with an assistant. The question is whether your collection is something that assistant can consult properly, or something it will guess about. The answer is a small server that exposes your holdings as a set of defined operations: list the works, fetch a specific record, retrieve an exact passage, search the text, find related items.
The design decision that matters is that no language model sits inside it. It reads from your published structure and returns text and records. The reasoning happens in whatever assistant the researcher brought, which means every improvement in those tools lifts your archive without you paying for it or rebuilding anything.
Whatever they already use. You do not choose it, host it, or pay for it.
List, fetch, get an exact passage, search, find related. No model inside. It serves your structure and nothing else.
Material you trust may be quoted with an address. Material you do not may only be described, with a link out. The interface applies those rules itself, so they hold even when nobody is watching.
This is how the think tank archive is published, and it is open. An institution that does this stops being a place assistants hallucinate about and becomes a place they cite.
Saying What You Actually Trust
Not every holding is equally readable. A clean typescript and a damaged 1890 page do not deserve the same confidence, and publishing both as though they did is how an archive ends up quoted wrongly. So the collection declares, per item, what may be done with it.
| Full text published | Paragraph addresses | May be quoted directly | Must link to the original | |
|---|---|---|---|---|
| Trusted text | yes | yes | yes | optional |
| Scanned works | no | no | no | yes |
Scroll the table sideways →
The tier is a field on the record, so it travels. Anything reading the collection can tell what it is permitted to assert before it says anything. The route from the lower tier to the upper one is already in the schema, so cleaning up a scan later changes the data rather than the design.
Knowing How You Are Cited
Most archives cannot answer a simple question from their own board: who used us this year, and for what. Reader registers record visits, not influence. So we build the citation picture from the outside in.
| Question | How it gets answered |
|---|---|
| Who cites us | Scholarly databases and open citation indexes are matched against your holdings, so a work in your collection carries the papers that cite it. |
| Which holdings matter | Citation counts attach to items, so acquisition and conservation budgets can follow demonstrated use rather than intuition. |
| Are we cited correctly | Malformed and broken references to your material are found and reported, and stable addresses are published so future citations resolve. |
| Where are we invisible | Fields that should be citing you and do not, which is a collections and outreach finding rather than a technical one. |
| Who uses us without saying so | Passages from your holdings appearing in published work without attribution, found by matching text rather than by trusting a bibliography. |
The Physical Layer, Tracked
Digitising a document does not retire it. Boxes still move, tapes still degrade, and a loan still has to come back. So we model the physical item alongside the digital one, and keep the two joined.
- Where it is now, and where it was
- Condition at last inspection
- Carrier, and whether that carrier is obsolete
- Loan and retention clocks
- Which object it was made from
- When, and at what settings
- The original bytes, kept unaltered
- Every later version derived from those
Keep the two joined and the institution can answer a question most catalogues cannot: which holdings sit on a format we will soon have no machine to read.
Carriers heading for obsolescence are then ranked by risk, rather than discovered on the day nobody can find a machine to read them.
Turning the Backlist Into Something Readable
Institutions sit on out-of-print books, journal runs and reports whose only form is a PDF nobody reads on a phone. Once the text is structured, the same source can be issued as a web edition with stable addresses for citation, as a reflowable electronic book, and as a fresh print-ready file. One structured source, several editions, rather than three separate retyping projects.
