Use cases

Collections That Answer Questions

An archive already holds the knowledge. What it lacks is a way for that knowledge to be asked a question, by a researcher or by anything else.

We have done work of this kind. See the think tank archive for how the mapping works in practice, and the open library at Falsafa.ai for the method in public.

The Path a Collection Takes

1Capture

Scanning, photography of bound and fragile material, and recording. What exists only on paper or only in someone's memory becomes a file that can be worked with.

2Make readable

Pictures of pages become text, including old typesetting, many scripts and handwriting. Recordings become transcripts with speakers and timings.

3Structure

The ontology: what each item is, who made it, what it is about, who is named inside it, and how items relate. Every relation carries the passage that proves it.

4Publish

To people as a fast public platform, and to machines as an open interface. Both read the same structure.

Most digitisation stops after step 2. That produces a searchable pile, which is better than a cupboard and still not a knowledge system. Steps 3 and 4 are where a collection becomes something you can ask a real question of.

Oral History as a First-Class Source

Testimony is the most perishable holding any institution has, and the least well handled. We record it, transcribe it with speakers separated and timings kept, and translate where the interview is in a language the catalogue is not. Mixed-language speech is tagged honestly as mixed rather than forced into one language, because that tag decides how strictly a later quotation can be trusted.

The transcript is then treated exactly like a manuscript. People and places named in it become entities. Passages get a fixed address so they can be cited. What the speaker said about a subject becomes a relation, with the words attached.

Material the Machine Finds Hard

Older holdings break ordinary pipelines: unusual scripts, marginalia, damaged pages, mixed languages on one leaf, and layouts no modern reader expects. Two commitments make this survivable. The original stays untouched and byte-exact, so every later pass can be rerun from the source rather than from a processed copy. And the catalogue declares how confident it is in each reading rather than presenting a shaky transcription as fact.

An Open Interface, on Your Terms

Researchers now arrive with an assistant. The question is whether your collection is something that assistant can consult properly, or something it will guess about. The answer is a small server that exposes your holdings as a set of defined operations: list the works, fetch a specific record, retrieve an exact passage, search the text, find related items.

The design decision that matters is that no language model sits inside it. It reads from your published structure and returns text and records. The reasoning happens in whatever assistant the researcher brought, which means every improvement in those tools lifts your archive without you paying for it or rebuilding anything.

The researcher bringsTheir own assistant

Whatever they already use. You do not choose it, host it, or pay for it.

You publishA defined set of operations

List, fetch, get an exact passage, search, find related. No model inside. It serves your structure and nothing else.

Your rules travel with itTiers and citation duties

Material you trust may be quoted with an address. Material you do not may only be described, with a link out. The interface applies those rules itself, so they hold even when nobody is watching.

This is how the think tank archive is published, and it is open. An institution that does this stops being a place assistants hallucinate about and becomes a place they cite.

Saying What You Actually Trust

Not every holding is equally readable. A clean typescript and a damaged 1890 page do not deserve the same confidence, and publishing both as though they did is how an archive ends up quoted wrongly. So the collection declares, per item, what may be done with it.

Full text publishedParagraph addressesMay be quoted directlyMust link to the original
Trusted textyesyesyesoptional
Scanned worksnononoyes

Scroll the table sideways →

The tier is a field on the record, so it travels. Anything reading the collection can tell what it is permitted to assert before it says anything. The route from the lower tier to the upper one is already in the schema, so cleaning up a scan later changes the data rather than the design.

Knowing How You Are Cited

Most archives cannot answer a simple question from their own board: who used us this year, and for what. Reader registers record visits, not influence. So we build the citation picture from the outside in.

QuestionHow it gets answered
Who cites usScholarly databases and open citation indexes are matched against your holdings, so a work in your collection carries the papers that cite it.
Which holdings matterCitation counts attach to items, so acquisition and conservation budgets can follow demonstrated use rather than intuition.
Are we cited correctlyMalformed and broken references to your material are found and reported, and stable addresses are published so future citations resolve.
Where are we invisibleFields that should be citing you and do not, which is a collections and outreach finding rather than a technical one.
Who uses us without saying soPassages from your holdings appearing in published work without attribution, found by matching text rather than by trusting a bibliography.

The Physical Layer, Tracked

Digitising a document does not retire it. Boxes still move, tapes still degrade, and a loan still has to come back. So we model the physical item alongside the digital one, and keep the two joined.

The physical itemStill exists, still moves
  • Where it is now, and where it was
  • Condition at last inspection
  • Carrier, and whether that carrier is obsolete
  • Loan and retention clocks
The digital surrogateKnows what it came from
  • Which object it was made from
  • When, and at what settings
  • The original bytes, kept unaltered
  • Every later version derived from those

Keep the two joined and the institution can answer a question most catalogues cannot: which holdings sit on a format we will soon have no machine to read.

Carriers heading for obsolescence are then ranked by risk, rather than discovered on the day nobody can find a machine to read them.

Turning the Backlist Into Something Readable

Institutions sit on out-of-print books, journal runs and reports whose only form is a PDF nobody reads on a phone. Once the text is structured, the same source can be issued as a web edition with stable addresses for citation, as a reflowable electronic book, and as a fresh print-ready file. One structured source, several editions, rather than three separate retyping projects.

Tell us your hardest problem. We will solve it.