Graphic Stories at Scale, Every Line Traced to a Source
One of India's largest publishers wanted graphic stories faster than any studio could make them, and accurate enough to put a living person's words in a speech bubble. Provenance became a mechanical gate rather than an editorial promise.
The pipeline is deployed and running. Titles come off it continuously across five publishing lines, and each new line we add extends the same machinery instead of needing its own.
The Ontology, Top to Bottom
Everything in the system is one of five things, and every line the pipeline produces knows where it sits. A new product line can then be added without touching the machinery.
Pace and accuracy pull against each other, and provenance is an editorial promise somebody has to keep by hand.
Provenance is a mechanical gate. Every beat points at a source line, and the gate resolves each pointer before the script proceeds.
A product family, and they are not alike. Illustrated biographies. India's classical epics and traditional knowledge. Health and social-awareness titles for children and teenagers. An original character universe for early readers. Activity and early-learning books. Each has its own format, length and voice.
A series inside the line. Business figures. Science figures. One epic. One awareness category such as road safety or mental health.
The person, character or topic a book is about. One subject can spawn several books, an early-years title and a later-career title, or one story told for two different age groups.
The research. Sources, each carrying its own provenance record, and an index regenerated from the files so it can never drift from what is actually there.
One script file per title, with its own metadata at the top, in a grammar a parser reads. Length, format and narrator are declared here, per title.
The contract is uniform even where the shape is not. A dossier built from converted books looks different inside from one built from authored topic notes, but both present the same three things to the tooling. That is why an entire new line was added by extending one list of permitted values, inheriting the grammar, the gates and the deliverables unchanged.
How the Research Gets In
Books, long interviews and video all arrive as something a machine cannot read straight: a scanned page, an audio file, a video. Each is converted to text, then split so that it can be pointed at precisely.
Books to text. Audio and video to transcripts, marked with who is speaking and when, and tagged honestly by language including mixed-language speech.
One file per chapter, under a chapter map that summarises each in a line. Only then can a claim point at an exact line.
Per source: where it came from, who made it, what year, and the licence position.
Generated from the files themselves, never hand-maintained, so the index cannot describe a library that no longer exists.
Chapter-per-file is what makes the rest possible. Because a claim can point at a file and a line number, a script can check the pointer. Without that granularity, "cite your source" stays a slogan instead of becoming a gate.
What the Subject's Own Books Never Mention
This is the part that surprises people. An authorised biography is a careful document. The vivid, candid material about someone is usually in other people's accounts: a rival describing them, a colleague recalling a bad decision, a mentor talking about them at nineteen.
Because every subject's research sits in one corpus rather than in a separate folder, the library holds those passages already. The work is finding them, and a plain text search is not good enough to do it.
Transcripts mangle surnames. Speakers refer to people by their company, their team, their role, or the one event everyone remembers, and never by name. Speaker labels are sometimes simply wrong.
Search on the name, its likely mangled forms, and the companies, teams and signature events attached to the person. Then read the passages rather than pattern-match them.
Only the person's own words count as that person speaking. Who is talking is decided from what is said, not from the label on the line. Each find is stored with its exact location and a confidence mark.
The first full sweep proved the point. It surfaced a strong passage that a plain text search had missed completely, because the subject's name appeared only in the interviewer's question and never in the answer that mattered.
What the Gate Caught
Every line of dialogue, narration and caption in a script carries a pointer to the source line behind it. Before a script moves on, a checker resolves every pointer against the actual file. If a pointer does not resolve, the script does not proceed.
- Read every pointer in the script
- Open the file and the line
- Confirm the claim is supported there
- Refuse the script if any pointer fails
- A draft with no pointers at all
- passes, because there is nothing to fail
- It reports success
- while guarding nothing
This actually happened, and we learned more from it than from anything that worked. One draft arrived with no citations in it at all. The gate passed it, truthfully and uselessly, because there was nothing to check. The checker now counts how much of the script is cited as well as checking the citations it finds.
The checker opens each one against the real file in the research corpus.
To art direction, then to the editorial review application, where the source link stays live on every beat so a reviewer can check the same thing by hand.
It is returned. There is no option to record the failure and continue, because a warning nobody has to clear is a warning nobody clears.
The gate is binary on purpose. Most quality tooling produces a report, and reports get skimmed under deadline. A stage that simply refuses to hand the work onward cannot be skimmed.
Preparation Is Not a Scratchpad
The quote bank assembled during research is an intermediate file, so it was treated more casually than a script. Wording that had been smoothed while making notes was later inherited into finished scripts as though it were sourced material, and one subject's words had to be un-quoted after the fact.
The rule that came out of it: a preparation artefact carries the same truth burden as the final page, because the final page will trust it without asking.
A Declared Length Is a Creative Device
Lengths differ by line and by title. What does not differ is that each title declares its length up front, and the machine then enforces that the script contains exactly that many pages, no more and no fewer.
A whole life, or a whole epic, then has to fit a canvas that cannot stretch, and that constraint is what forces the narrator's voice into a definite shape. Declaring the length per title is also what lets a 48-page biography and a short awareness title run through one pipeline without either being bent into the other's shape.
The Art Is Generated Against a Locked Look
The pipeline does not stop at a script for someone else to draw. Frontier image models generate the characters, the backgrounds and the finished pages. The hard part is not making one good picture; it is making the four hundredth picture of the same character still look like that character.
So every recurring figure carries a locked visual specification with reference art, and a figure belongs to exactly one place that owns its look even when the character appears across several programs. Props and settings are locked the same way. Generation happens against those locks rather than against a fresh description each time.
One deliberate exception: for gods and legendary figures, what a reader already recognises outranks what a text technically describes. A textually defensible figure a child cannot recognise is a failed design. The sourcing contract still governs every word they say and everything they do.
What Comes Out
One file, fixed grammar, declared length, every beat sourced.
Characters, backgrounds and finished pages generated against the locked specifications.
A private application where the publisher's team reads each title with the source link live on every beat.
Print-ready colour files, editable text versions, and translated editions.
Everything reads the same structure. The validator, the renderer and the citation checker all use one parser, so an editor can never be shown a page different from the one that was checked.
Most generated-art pipelines fail at the last box, because the tools stop once there is a picture. A printer does not want a picture. A printer wants colour separated for the press it is actually running, with black text that is genuinely one ink rather than a muddy mix of four, at a resolution the paper can hold.
An editor, meanwhile, wants to change a word without regenerating a page. An image model produces neither, so we built both. One converts the colour so it survives uncoated paper. The other turns a rendered page back into a document you can edit, in the same lettering as the printed book.
Separated for the press, with text held to a single black ink so it stays crisp instead of blurring across plates, at print resolution throughout.
Rendered pages converted back into an editable document in the book's own lettering, so a correction is a text edit rather than a regeneration.
This is the dull half of the work, and it is the half that decides whether a book gets printed. Generating a good-looking page is the easier part. Getting that page through a press and past an editor is what turns it into a book.
