All writing

The Author Can Only Do Half

Diagram: an author binds notes, headers and captions inside a tagged PDF; the pipeline must then write that structure beside each piece of evidence the model reads

This closes a three-paper cycle on retrieval over tagged documents. Evidence Graph Retrieval argued that a tagged PDF should be compiled, not chunked. Document Task Answerability let the document write its own exam, and found that files scoring 98–100 on every accessibility checker were still only 42% answerable through the best retriever. That paper ended with eighteen authoring guidelines — and a guideline derived from a failure profile is a hypothesis, not a lever, until someone edits a document and the score moves. Part 3 is that test.

The same eight PDFs were re-tagged by hand, by the person who tagged them the first time, along those guidelines: notes bound explicitly to the claims they qualify, captions and sectioning, per-cell header pointers, a clean text layer, and an abbreviation list plus a small block of professional knowledge attached as metadata. Nothing a sighted reader sees changed. The benchmark side was frozen — tasks, gold and the four validated phrasings reused byte-for-byte — so every difference is attributable to the intervention. (A naming note: the structure-preserving retriever called EGR-S in the first two papers is written Egres from here on, pronounced “egress”. Same system.)

What moved, and what did not:

  • Half the pilot’s retrieval failures were never retrieval failures. The evidence was found but ranked below what the model was shown. Widening the evidence window on the original documents took Egres from 42% to 54% before a single tag was touched.
  • Re-tagging what was already tagged changes almost nothing — 54% to 55%. A retriever that reads the tag tree already had the headings, tables and lists; wrapping them in more containers gives it nothing new to walk. One perfectly conformant edit, a caption in place of a heading, cost a document nine points until the retriever learned to read captions as titles.
  • The real levers are declarative. Binding footnotes to their claims and pointing each data cell at its headers changed what the document can state: the tax guide went from 2 to 18 bindable qualification tasks. That shows up as a larger task universe and unambiguous provenance, not as a higher score on the old tasks — which is why an intervention study must re-enumerate, not just re-score.
  • The gain appears only when the pipeline writes the structure next to the evidence. The answering model had been receiving each cell as [id] (TABLE_CELL p2) 16.2 — no row header, no column header, no caption, no note. Serializing those beside each value lifted Egres from 56% to 59% and cleared answer-stage failures no document edit had been able to touch. The chunk baseline, with no structure to carry, did not move.
  • Attached knowledge helps exactly where its terms occur, and nowhere else. Abbreviations and domain definitions as metadata were worth +2.4 points on tasks using a glossed term, and cost a little on unrelated tasks when applied globally. Gloss per item, never as a preamble.

The conclusion is that answerability is a property of the pair — document and consumer — and the two halves are not symmetric. The author can declare relations the document could not previously state, keep the text layer honest, and attach the knowledge a reader is assumed to bring. All of that is worth nothing until a consumer carries it through to the model. The paper ends with a matched pair of lists: thirteen rules for taggers whose payoff depends on the pipeline, and five for the people building the pipeline. Between them, a good part of the distance between compliant and answerable can be closed from both ends.

The full paper — intervention, re-assessment protocol, both rounds of results and the guidance — is here: Making a Tagged PDF Answerable — Cyphia Solutions (2026), Part 3 of Document Task Answerability.

Read the full paper

The intervention, the re-assessment protocol, both rounds of results and the guidance for taggers and pipeline builders are in Making a Tagged PDF Answerable.