← Blog Cyphia Solutions LLC
Research paper · Part 3 · intervention study · draft v0.2

Making a Tagged PDF Answerable: What Remediation and Metadata Do — and Do Not — Change

Part 2 measured how much of a document's own task universe a RAG system can complete. Part 3 intervenes on the documents: the eight pilot PDFs were remediated by hand along the Part 2 guidelines (explicit note binding, captions and sectioning, key–value tables as lists, clean text layer), re-compiled into evidence graphs, and re-assessed on the identical frozen tasks and questions. The result separates what document structure can fix from what the retrieval pipeline must fix — and finds that the semantic edits only pay off once the pipeline writes structure into what the model sees.

Cyphia Solutions LLC · 28 August 2026 · Part 3 of Document Task Answerability (Cyphia Solutions, 2026) · Builds on Evidence Graph Retrieval (Cyphia Solutions, 2026)

Nomenclature: the structure-preserving retriever written EGR-S in Parts 1 and 2 is written Egres here (pronounced “egress”); the system is unchanged.

Abstract

Part 2 showed that PDFs which pass every accessibility checker are still far from answerable through retrieval-augmented generation. Part 3 asks what to do about it, by intervening on both sides of the gap at once: the eight pilot documents were re-tagged by hand along the Part 2 guidelines — explicit note binding, captions and sectioning, per-cell header pointers, a clean text layer, and an abbreviation list plus a small block of professional knowledge attached as metadata — and the retrieval pipeline was changed so that what a document declares is actually carried to the language model. The frozen tasks, gold and phrasings of Part 2 were reused unchanged, so every difference is attributable to the intervention.

Three observations organise the results. First, structure a retriever already had cannot be improved for it: re-tagging what was already tagged left evidence reach unchanged, and one conformant edit — a caption in place of a heading — briefly made a document worse until the retriever learned to read captions as titles. Second, the levers are declarative: binding notes to the claims they qualify and pointing each data cell at its headers change what the document can state, which shows up as a larger task universe and as unambiguous provenance rather than as a higher score on the old tasks. Third, and decisively, the gain appears only when the pipeline writes the structure next to the evidence: with headers, captions, headings and bound notes serialized beside each value, answer-stage failures fall sharply for the structure-aware system while the structure-blind baseline does not move, and injected domain knowledge helps exactly where its terms occur and nowhere else. From these we propose a short set of authoring guidelines for taggers — header pointers, explicit bindings, additive captions, scoped metadata — that promise to make answerability an attainable property of a tagged PDF, provided consumers carry the structure through.

42% → 54%
Egres, same documents: a wider evidence window (8,000 chars → 1,400 words). Half of the pilot's retrieval failures were evidence retrieved but not shown.
54% → 55%
Egres, remediated documents at the same window: re-tagging what was already tagged changes almost nothing.
56% → 59%
Egres, headers · captions · headings · bound notes written beside each value: the structure finally reaches the model.
+2.4 pts
Egres, abbreviations and professional knowledge attached as metadata — on the tasks that use a glossed term; nothing elsewhere.
30% → 48% → 48%
chunk baseline: moves only with the window, flat under remediation, context and knowledge — it has no structure to carry.
2 → 18
qualification tasks the tax guide can generate once notes are bound explicitly — a larger universe, not a higher score on the old tasks.

1 Why intervene

Part 2 left a specific question open. Eight documents that pass accessibility checkers at 98–100 yielded a task universe of which the best retriever completed 42% of task–phrasing pairs. The paper's §8 turned the diagnosis into eighteen authoring guidelines, but a guideline derived from a failure profile is a hypothesis, not a lever, until a document is changed and the score moves. Part 3 is that test: the same eight PDFs, edited by the same person who tagged them, along the guidelines, with the benchmark side frozen — tasks, gold and the four validated phrasings are reused byte-for-byte.

The failure analysis that preceded the intervention (on the 5,175 pilot runs) also changed what we expected to move. Only 22% of tasks were unanswerable by every system and phrasing; 45 of those 97 had full evidence in front of the model and failed at the answer stage. Half of the retrieval-class failures were context-window losses — the evidence was retrieved but ranked below the shown prefix — not reach failures. And phrasing was the largest single factor (0.55 → 0.26 across the four phrasings of the same task). So the pre-registered prediction (H4b) was that gains, if any, would concentrate in retrieval-class failures on the documents whose punch lists were worked, while the answer stage and lexically distant phrasings would move little.

2 The intervention

2.1 Edits made to the documents

Remediation was done in the tagging tool under a content-integrity rule: every gold value, list item and note text remains in the document unchanged; headings are not renamed; allowed edits are the guidelines' own — tags, scopes, sectioning, captions, alt text, ActualText, note and link binding. The compiled-graph diff (Table 1) is the authoritative record of what changed.

Table 1 — Structural change per document (original → improved; unchanged counts shown once).
documentppTableList ItemCaptionSectionHeadingNoteFigurefigures w/ altHEADING_FORCAPTION_FORREFERENCES
D1 FixedIncomeFund fact sheet2101000→1017020→2101→4700
D2 CorePlusBondFund fact sheet414→13100→60→1628→22166192→1380→60
D3 2025 Tax Guide526→2123→440038→403422250→25800→58
D4 Ed 43 Summary of Changes17360001801121000
D5 Long/Short Growth Equity exposure report41711002373313700→7
D6 PIMCO VIT CommodityRealReturn QIR131135003877→67→6341→33900→12
D7 Course Outline12954004501128200
D8 Hilton FY2024 Slavery & Trafficking Statement157→444→4910190551950→10

The edits cluster into four kinds. Explicit note binding (D3, D5, D6): footnotes re-tagged as FENote with reference pointers that the updated dump exports as refer_tag; 58 bindings in the tax guide alone. Captions and sectioning (D1, D2): table blocks wrapped in Sect with a Caption, and — this matters below — in D2 the caption replaced the heading (headings 28→22). Key–value tables as lists (D3, D8): two-column label/value tables converted to L/LI/Lbl/LBody. Text-layer hygiene (D1, D8): kerning artefacts ("F un d") fixed; composite cells given ActualText. D4 and D7 show no tag-tree change and act as controls.

2.2 Pipeline changes forced by the edits

Three of the edits exposed gaps in the consumer, which were fixed so that the intervention is measured against a retriever that understands the structures the guidelines ask for. (a) FENote was unmapped (34 notes in D3 silently became untyped); it now maps to NOTE, and refer_tag compiles to REFERENCES edges (referrer→note, direction-normalised). Expansion follows them both ways, and the NOTE_QUALIFY motif binds by explicit reference first, marker match second. (b) Captions: the compiler now emits CAPTION_FOR (caption→captioned table/figure/list) and, when a container opens with a caption instead of a heading, caption-as-title scope; captions join the structural seed channel and the index's context projection; a caption seed expands into its table. (c) Gold re-anchoring: node ids change on re-tagging, so a re-anchorer maps each gold reference by text (exact → whitespace-squashed → thousands-separator-insensitive → substring → alt text → role/page/ordinal), disambiguated by header/label/heading context, keeping both the original and the new text-layer string as accepted evidence. All 436 tasks re-anchored with no unresolved reference.

3 Re-assessment protocol

Two layers, because they answer different questions. The retrieval layer is deterministic: for every paired run, evidence recall and relation coverage are recomputed from the stored retrieval traces on the full retrieved list (reach) and on the 1,400-word shown prefix. No language model is involved, so the paired delta is exactly the effect of the document edits plus the consumer fixes above. The answer layer repeats Part 2's scoring — answerer over the shown evidence, judge over proposition checklists, strict PASS contract — but the pilot had run with an 8,000-character window (one chunk shown) while the current protocol shows 1,400 words (two chunks). A first pass against the pilot scores produced a +6.4-point "gain" for chunking on byte-identical evidence; the original substrate was therefore re-answered and re-judged at 1,400 words so that before and after share one protocol. That re-baseline is the comparison reported here; the pilot-protocol comparison is kept only as a measure of answerer/judge drift.

Statistics: exact McNemar on discordant pairs per system; a task-clustered bootstrap CI on the mean ΔPASS; per-document McNemar for the document table. Pre-registered hypotheses are in docs/dtaa-phase3-plan.md.

4 Results

4.1 Retrieval layer (deterministic)

Table 2 — Retrieval-level paired deltas, all documents, all phrasings (5,175 pairs per system across 3 systems). Reach = full retrieved list; shown = 1,400-word prefix.
systemevidence recall (reach)full evidence (reach)relation coverage (reach)evidence recall (shown)full evidence (shown)
chunk hybrid0.823→0.827 (+0.004)0.656→0.667 [0→1 20, 1→0 0, p=1.9e-06]0.768→0.7800.608→0.6120.405→0.411
Egres0.846→0.846 (+0.000)0.699→0.701 [0→1 16, 1→0 13, p=0.71]0.760→0.7580.806→0.8080.646→0.650
Egres (reconstructed)0.600→0.604 (+0.004)0.344→0.354 [0→1 17, 1→0 0, p=1.5e-05]0.485→0.4930.539→0.5430.299→0.309
Table 3 — Evidence reach per document (before→after, Δ points).
documentchunk hybridEgresEgres (reconstructed)
D10.94→0.98 (+4)0.88→0.89 (+1)0.58→0.61 (+4)
D20.85→0.85 (+0)0.86→0.87 (+1)0.50→0.50 (+0)
D30.76→0.76 (+0)0.93→0.93 (-0)0.73→0.73 (+0)
D40.96→0.96 (+0)0.91→0.91 (+0)0.73→0.73 (+0)
D50.91→0.91 (+0)0.67→0.67 (+0)0.58→0.58 (+0)
D60.82→0.82 (+0)0.73→0.73 (+0)0.51→0.51 (+0)
D70.67→0.67 (+0)0.83→0.83 (+0)0.55→0.55 (+0)
D80.79→0.79 (+0)0.95→0.93 (-2)0.67→0.67 (+0)

The intervention is neutral at the retrieval layer for the structure-preserving retriever and for the reconstructed substrate; the structure-blind chunker gains where the text layer was cleaned (D1: kerning). The only per-document movement of size was D2, where turning headings into captions removed the tables from their section scope: Egres reach fell nine points on that document until captions were compiled and consumed as titles, after which it is one point up. Every motif-level Δ lies within ±3 points.

4.2 Answer layer (like-for-like 1,400-word budget)

Table 4 — Task answerability before→after, same tasks, same questions, same budget.
systemrunsAt beforeAt afterΔ pts0→11→0McNemar p95% CIevidence recallwords shown
chunk hybrid17250.3580.364+0.61000.002[+0.1, +1.3]0.61→0.611162→1162
Egres17250.5440.547+0.228240.68[-0.7, +1.3]0.81→0.811042→1041
Egres (reconstructed)17250.2610.271+1.01701.5e-05[+0.3, +1.9]0.54→0.541131→1131
Table 5 — Per document (At before→after, Δ points; * McNemar p < 0.05 within the document).
documentchunk hybridEgresEgres (reconstructed)
D1 FixedIncomeFund fact sheet0.55→0.60 (+5*)0.64→0.63 (-1)0.26→0.35 (+9*)
D2 CorePlusBondFund fact sheet0.51→0.51 (+0)0.63→0.70 (+7*)0.29→0.29 (+0)
D3 2025 Tax Guide0.27→0.27 (+0)0.62→0.60 (-2)0.38→0.38 (+0)
D4 Ed 43 Summary of Changes0.39→0.39 (+0)0.65→0.65 (+0)0.37→0.37 (+0)
D5 Long/Short Growth Equity exposure report0.27→0.27 (+0)0.34→0.34 (+0)0.15→0.15 (+0)
D6 PIMCO VIT CommodityRealReturn QIR0.31→0.31 (+0)0.34→0.34 (+1)0.13→0.13 (+0)
D7 Course Outline0.35→0.35 (+0)0.61→0.61 (+0)0.33→0.33 (+0)
D8 Hilton FY2024 Slavery & Trafficking Statement0.25→0.25 (+0)0.48→0.45 (-3*)0.14→0.14 (+0)
Table 6 — Where the movement is: pass rate after, by the run's failure class before (H4b).
failure class at baselinechunk hybrid n / before→after (Δ)Egres n / before→after (Δ)Egres (reconstructed) n / before→after (Δ)
R1005 / 0.00→0.01 (+1)596 / 0.00→0.02 (+2)1192 / 0.00→0.01 (+1)
A77 / 0.00→0.00 (+0)138 / 0.00→0.09 (+9)61 / 0.00→0.00 (+0)
G1 / 0.00→0.00 (+0)30 / 0.00→0.00 (+0)3 / 0.00→0.00 (+0)
Q24 / 0.00→0.00 (+0)22 / 0.00→0.09 (+9)18 / 0.00→0.00 (+0)
PASS618 / 1.00→1.00 (+0)939 / 1.00→0.97 (-3)451 / 1.00→1.00 (+0)
Table 7 — By phrasing (Q2 = lexically distant).
variantchunk hybridEgresEgres (reconstructed)
Q0433 / 0.42→0.43 (+1)433 / 0.64→0.64 (-0)433 / 0.30→0.32 (+1)
Q1433 / 0.38→0.39 (+1)433 / 0.59→0.60 (+1)433 / 0.27→0.29 (+1)
Q2429 / 0.28→0.28 (+0)429 / 0.38→0.38 (-1)429 / 0.18→0.18 (+0)
Q3430 / 0.35→0.35 (+0)430 / 0.57→0.58 (+1)430 / 0.29→0.30 (+1)
Table 8 — By structural motif.
motifchunk hybridEgresEgres (reconstructed)
FIGURE_ALT60 / 0.00→0.00 (+0)60 / 0.50→0.50 (+0)60 / 0.00→0.00 (+0)
HEADING_LIST76 / 0.49→0.49 (+0)76 / 0.63→0.63 (+0)76 / 0.33→0.33 (+0)
HEADING_SCOPE213 / 0.54→0.54 (+0)213 / 0.78→0.77 (-1)213 / 0.60→0.60 (+0)
LABEL_VALUE165 / 0.56→0.56 (+0)165 / 0.62→0.59 (-3)165 / 0.30→0.30 (+0)
NOTE_QUALIFY36 / 0.22→0.22 (+0)36 / 0.44→0.42 (-3)36 / 0.06→0.06 (+0)
SECTION_SET111 / 0.32→0.32 (+0)111 / 0.56→0.59 (+3)111 / 0.21→0.21 (+0)
TABLE_CELL360 / 0.29→0.30 (+1)360 / 0.49→0.49 (-0)360 / 0.22→0.24 (+2)
TABLE_COMPARE252 / 0.33→0.36 (+2)252 / 0.60→0.62 (+1)252 / 0.21→0.23 (+1)
TABLE_SERIES452 / 0.32→0.32 (+0)452 / 0.41→0.43 (+2)452 / 0.20→0.21 (+1)
Table 4b — The context budget moved answerability far more than the remediation did (same documents, same tasks, same questions).
systemAt pilot (8,000 chars)At original @1,400 wordsΔ budget (pts)At improved @1,400 wordsΔ remediation (pts)
chunk hybrid0.300.36+5.80.36+0.6
Egres0.420.54+12.40.55+0.2
Egres (reconstructed)0.230.26+3.10.27+1.0

Table 4b is the study's least expected result. Re-answering the original documents under the 1,400-word window instead of the pilot's 8,000-character one raised Egres from 0.42 to 0.54 and chunking from 0.30 to 0.36; the remediation then added +0.2 and +0.6 points respectively. Half of Egres' retrieval-class failures in the pilot were evidence that had been retrieved but not shown; widening the window is what converted them.

Reading Tables 4–8 against Table 2: whatever moves at the answer layer moved without the evidence set changing, so the chunk-hybrid column is the floor for answerer/judge variance under this protocol, and the Egres column should be read net of it (H4c: Egres − chunk ≈ -0.3 points). The retrieved-but-not-shown share of Egres' class-R failures is 135/596 before and 123/591 after — the budget, not the document, decides those.

4.3 The task universe after remediation (H4d)

Table 9 — Re-enumerating the universe on the improved graphs (no LLM).
documenteligible tasksAMBIGUOUS rejectionsNOTE_QUALIFY tasks (explicitly bound)
D1291→2930→00→0 (0 explicit)
D2344→3450→00→0 (0 explicit)
D3270→2790→02→18 (18 explicit)
D4889→8890→00→0 (0 explicit)
D52,068→2,0702→24→6 (4 explicit)
D6713→714131→1313→3 (3 explicit)
D784→850→00→0 (0 explicit)
D852→4910→100→0 (0 explicit)

Explicit binding is the one guideline whose effect is visible at the universe level: the tax guide's deterministically bindable qualification tasks go from 2 to 18 because the repeated */*** markers that defeated the page-unique rule are now resolved by the document itself. The ambiguity lint did not move (D6: 131 repeated header–value pairs; D8: 10) — captions and summaries were added to D2, where nothing was flagged, not to the documents where the lint fired; and converting D8's tables to lists removed three LABEL_VALUE tasks without disambiguating the flagged pairs.

5 What the intervention reveals

5.1 Structure that a retriever already had cannot be "improved" for it

Egres reads the tag tree. Adding Sect wrappers, captions and list tags to a document whose headings, tables and lists were already tagged gives it nothing new to walk; reach stays at 0.846. The edits change how the document describes its parts, not how they are connected — and description is consumed at the answer stage and in the index, neither of which was reading it (§5.4).

5.2 A guideline can backfire against a naive consumer

Replacing a heading with a caption is the correct PDF/UA form for a titled table block, and D2 lost nine points of reach for it. The lesson generalises: a structure standard's preferred encoding is only as good as the least capable consumer, and the guideline should say "add a caption; keep the heading" until consumers treat captions as titles. The consumer fix is small (CAPTION_FOR edges, caption seeds, caption scope) and should be part of any structure-aware retriever.

5.3 Explicit binding is the deterministic lever

The only edit that changed what the document can declare is note binding. Its effect is invisible on the frozen tasks — they were sampled from the universe the old document could generate, so they were bindable already — and visible on the universe the new document generates (2→18). The evaluation lesson is that an intervention study must re-enumerate, not only re-score: frozen tasks measure regression, the universe measures capability.

5.4 The answer stage never saw the structure

The answerer receives each evidence item as [id] (TABLE_CELL p2) 16.2: no row or column header, no caption, no heading path, no bound note. The wrong-neighbour reads that make up 135 of Egres' 200 answer-stage failures are the direct consequence, and no document edit can reach them. Likewise the index embeds ActualText | text | Alt only, so expansion text (E), table Summary and abbreviation lists — the guidelines' vehicles for phrasing robustness — never affect a seed. The remediation was aimed at the two failure classes the pipeline was structurally unable to pass on to the model.

6 Round 2: metadata that reaches the model

Round 2 acts on §5.4. On the tagging-tool side, every TD now carries explicit Headers pointers (compiled at top priority, origin SOURCE_ATTRIBUTE: D4 2,154/2,154 header edges declared, D1 387/387 — against 261 declared + 53 positional in the original D1), and an abbreviation list plus a block of reliable professional knowledge (the credit-rating hierarchy AAA … CAA/CCC as used by Moody's, S&P and Fitch) is attached to each document's metadata — no PDF structure involved. On the pipeline side, three arms on the improved substrate share documents, frozen tasks, questions, retrieval traces and an 1,800-word budget and differ only in what is written next to each evidence item: A bare item text; B + structural context from compiled edges (heading path, caption, row/column headers, list label, bound notes); C + abbreviation expansions and knowledge glossed onto the items that contain the term, and embedded into the index projection. The chunk baseline has no edges and receives no glosses, so it is the control for both factors.

At the retrieval layer the declared headers change nothing (Egres reach 0.846 in both) — the compiler was already deriving the same relations from Scope — so every effect below is an answer-layer effect.

Table 10 — Arms (all documents, all phrasings). Reference rows: original and improved documents at 1,400 words.
armchunk hybrid At / class-A share of failures / Q2Egres At / class-A share of failures / Q2Egres (reconstructed) At / class-A share of failures / Q2
orig@14000.358 / 0.07 / 0.280.544 / 0.18 / 0.380.261 / 0.05 / 0.18
impr@14000.364 / 0.07 / 0.280.547 / 0.17 / 0.380.271 / 0.05 / 0.18
A plain@18000.480 / 0.10 / 0.430.563 / 0.19 / 0.390.288 / 0.05 / 0.20
B +context0.478 / 0.11 / 0.420.595 / 0.13 / 0.420.292 / 0.04 / 0.20
C +context+knowledge0.478 / 0.11 / 0.420.598 / 0.15 / 0.420.291 / 0.05 / 0.20
Table 11 — Paired factor contrasts.
factorchunk hybrid Δ pts (0→1 / 1→0, McNemar p)Egres Δ pts (0→1 / 1→0, McNemar p)Egres (reconstructed) Δ pts (0→1 / 1→0, McNemar p)
budget 1,400 → 1,800 words (A − impr@1400)+11.6 (220 / 20, p=9.1e-44)+1.6 (35 / 7, p=1.5e-05)+1.6 (32 / 4, p=1.9e-06)
cell/heading/caption/note context (B − A)-0.2 (18 / 21, p=0.75)+3.2 (69 / 14, p=6.8e-10)+0.4 (26 / 19, p=0.37)
abbreviations + professional knowledge (C − B)+0.0 (0 / 0, p=1)+0.3 (43 / 38, p=0.66)-0.1 (11 / 13, p=0.84)
both (C − A)-0.2 (18 / 21, p=0.75)+3.5 (95 / 35, p=1.4e-07)+0.3 (25 / 19, p=0.45)

Table-cell context is the lever. Writing headers, captions, heading paths and bound notes next to the value moves Egres by +3.2 points (p = 7e-10) while the chunk baseline moves -0.2: 46% of Egres' answer-stage (class A) failures at arm A pass at arm B, notes +11 and section sets +10 points, the drug-formulary tables of D4 +15, and the lexically distant Q2 phrasing +3. This is the effect the round-1 remediation was aiming at and could not produce, because the structure never reached the model.

Table 12 — Injected knowledge, targeted: tasks whose intent or gold contains a glossed term versus the rest.
systemtasksrunsAt B → CΔ pts0→1 / 1→0McNemar p
chunk hybridtouching a glossed term7570.474 → 0.474+0.00 / 01
chunk hybridnot touching one9680.481 → 0.481+0.00 / 01
Egrestouching a glossed term7570.596 → 0.620+2.428 / 100.0051
Egresnot touching one9680.594 → 0.581-1.315 / 280.066
Egres (reconstructed)touching a glossed term7570.318 → 0.310-0.81 / 70.07
Egres (reconstructed)not touching one9680.272 → 0.276+0.410 / 60.45
Table 13 — Knowledge effect per document, tasks touching a glossed term.
document (terms)chunk hybrid n / B→CEgres n / B→CEgres (reconstructed) n / B→C
D130 / 1.00→1.0030 / 0.80→0.9730 / 0.40→0.37
D244 / 0.75→0.7544 / 0.70→0.7544 / 0.45→0.45
D355 / 0.07→0.0755 / 0.45→0.4955 / 0.05→0.05
D4149 / 0.66→0.66149 / 0.87→0.85149 / 0.49→0.47
D555 / 0.62→0.6255 / 0.29→0.4455 / 0.04→0.02
D669 / 0.33→0.3369 / 0.22→0.2369 / 0.17→0.16
D7335 / 0.38→0.38335 / 0.61→0.61335 / 0.36→0.36
D820 / 0.40→0.4020 / 0.35→0.4520 / 0.00→0.00

Professional knowledge works where it applies, and costs a little where it does not. Pooled over all runs the knowledge arm is flat (Egres +0.3). On the 757 Egres runs whose task involves a glossed term it is +2.4 points (p = 0.0051): the credit-quality tasks of D1 go from 0.80 to 0.97 once "BAA" is accompanied by "medium grade, lowest investment-grade tier", D5's exposure tasks from 0.29 to 0.44 (net/gross/market-cap definitions), D8 from 0.35 to 0.45. On the remaining runs it is -1.3 (p = 0.066): gloss lines spend budget. The chunk baseline is exactly 0 in both subsets (identical evidence, cached answers). Knowledge should therefore be attached per item, only where the term occurs — which is what the metadata-driven design does — not appended globally.

Task answerability A_t across the studyEgreschunk hybridEgres (reconstructed)0%20%40%60%Part 2 pilot8,000 charsoriginal docs@1,400 wordsremediated docs@1,400@1,800+ cell context+ knowledgeround 1 · documentsround 2 · what the model is shownEgres (reconstructed) · Part 2 pilot 8,000 chars: 0.230Egres (reconstructed) · original docs @1,400 words: 0.261Egres (reconstructed) · remediated docs @1,400: 0.271Egres (reconstructed) · @1,800: 0.288Egres (reconstructed) · + cell context: 0.292Egres (reconstructed) · + knowledge: 0.29123%29% · Egres (recon.) (+6 pts)chunk hybrid · Part 2 pilot 8,000 chars: 0.300chunk hybrid · original docs @1,400 words: 0.358chunk hybrid · remediated docs @1,400: 0.364chunk hybrid · @1,800: 0.480chunk hybrid · + cell context: 0.478chunk hybrid · + knowledge: 0.47830%48% · chunk hybrid (+18 pts)Egres · Part 2 pilot 8,000 chars: 0.420Egres · original docs @1,400 words: 0.544Egres · remediated docs @1,400: 0.547Egres · @1,800: 0.563Egres · + cell context: 0.595Egres · + knowledge: 0.59842%60% · Egres (+18 pts)
Figure 1 — Task answerability At (strict PASS, mean over four phrasings) at each stage of the study, same 436 tasks throughout. Egres: 42% in the Part 2 pilot → 60% with remediated documents, an 1,800-word window, serialized cell context and injected knowledge. The chunk baseline moves only with the window (30% → 48%) and not at all with context or knowledge — it has no structure to serialize. Reconstructed structure stays near 29%. Hover a point for its value; Tables 4, 10 and 11 give the paired statistics.

The budget, again. Raising the shared window from 1,400 to 1,800 words moves the chunk baseline +11.6 points (a third chunk fits) and Egres +1.6. Under the same budget Egres with context (0.59) leads chunking (0.48) by twelve points on the improved documents.

7 Practical guidance

What the two rounds support, stated as instructions. The evidence behind each item is the table it cites; items without a measured effect are marked as such.

7.1 For taggers: making a PDF answerable, not merely compliant

T1 — Give every data cell explicit header pointers. Set Headers on each TD (all row and column headers that apply, including group headers), not only Scope on the TH. Retrieval reach does not change (the compiler derives the same relations from Scope), but the declared pointers are what let a consumer print "row ‘Chile’ · column ‘Fund’" beside a value with certainty, and they resolve repeated values unambiguously (D2: 69 ambiguous gold references → 4). Round 2 rests on them.

T2 — Bind notes explicitly. Tag footnotes as FENote (or Note) with reference pointers from the marker to the note; never rely on a bare * being unique on the page. This is the one edit that expanded what the document can declare (D3: 2 → 18 bindable qualification tasks, Table 9) and it is what lets the consumer show "note: …" beside the claim (NOTE_QUALIFY +11 points in round 2).

T3 — Caption every table and figure, and keep the heading. A Caption as first child of the table block is right; a caption instead of a heading removed the table from its section scope for a heading-only consumer (D2 −9 points of reach, §5.2). Until every consumer treats captions as titles, add — do not replace.

T4 — Make the caption say what the table is. "Country distribution (%), fund vs. index, as of 31 Mar 2025" beats "Table 3". Captions are indexed and shown as context; a caption that repeats the document title adds nothing (Table 1's D2 captions did, and the gain came from the scope, not the words).

T5 — Put units and periods in headers, one value per cell. "$1,192.50 plus 12%" in one cell, or a unit that lives only in a footnote, defeats numeric matching and comparison. Composite cells were among the round-1 answer-stage residuals; splitting them is a source edit no pipeline can replicate.

T6 — Keep the text layer machine-clean. Fix kerning artefacts ("F un d"), keep ActualText consistent with the visible string (the "7,566" vs "7566" case broke string matching for every system), and avoid ActualText that changes the number's form. This was the only document edit that moved the structure-blind baseline (D1 +4 reach).

T7 — Disambiguate repeated header–value pairs. When the same header pair maps to different values in one section (D6: 131 cases, D8: 10), add a caption or a qualifying header ("2Q24", "Class Y"); the answerability lint lists them. No consumer fixed these and none can.

T8 — Attach an abbreviation list as document metadata. Every "BAA", "DIN/PIN", "MAGI", "PVITINST" the document uses, with its expansion — none of the eight files carried an E attribute. Round 2 shows expansions and short domain definitions attached per term are worth +2.4 points on the tasks that use them (Table 12); the same text placed nowhere is worth nothing.

T9 — Attach reliable professional knowledge, scoped to the terms that occur. Standard hierarchies (rating tiers, dosage forms, share-class codes) belong in metadata, not in the page. They helped exactly where the term appears (D1 credit quality 0.80 → 0.97) and cost budget where it does not (−1.3 on unrelated tasks); the consumer must apply them per item.

T10 — Key–value tables as lists only when the pairs are self-contained. Converting two-column label/value tables to L/LI/Lbl/LBody is legitimate, but it removed three tasks in D8 and cost 3 points there (Table 5): a list body loses the column header that told the reader what kind of value it is. If the table had meaningful column headers, keep it a table with T1.

T11 — Section headings that match the reading. An H nested under the wrong parent ("Key employee" under "Highly compensated employee") produces gold the reader disagrees with and answers the judge rejects. Levels should follow the visual hierarchy, and every Sect should open with its own heading or caption.

T12 — Informative alt text, or none. Figure tasks are answerable only through alt text (chunk baseline 0.00 on FIGURE_ALT); "Brandmark of …" answers nothing. State what the figure conveys, with its numbers.

T13 — Do not re-tag what already works. Wrapping already-tagged headings, tables and lists in more containers changed nothing for a structure-aware retriever (§5.1). Spend the effort on T1, T2, T5–T9.

7.2 For pipeline builders

P1 — Serialize structure into the evidence block. Heading path, caption, row/column headers, list label and bound notes next to each item: +3.2 points for Egres, none for chunks (Table 11). The cheapest change in the study and the largest.

P2 — Apply glosses per item, never globally. Expansions and knowledge only where the term occurs; a global preamble spends the window on unrelated tasks (−1.3 where the term is absent).

P3 — Budget per evidence bundle. A cell travels with its headers and caption, a claim with its note; trim bundles, not items. Report "retrieved but not shown" separately: it was half of the retrieval failures, and the 1,400 → 1,800-word step alone was worth +11.6 points to the chunk baseline.

P4 — Read captions as titles and follow reference edges. CAPTION_FOR and REFERENCES are small additions to a heading-only walker and they are what T2/T3 assume.

P5 — Treat declared relations as provenance. Keep the origin of every header edge (declared / scope-derived / positional); it tells the reader how much to trust a pairing.

7.3 For evaluators

E1 — Re-enumerate, not only re-score. Frozen tasks measure regression; the universe measures capability (T2's effect was invisible on frozen tasks).

E2 — Hold the protocol fixed and keep a control. Half the first before/after was a budget change; the structure-blind baseline on identical evidence is the noise floor.

E3 — Report At with its companions. Any-phrasing answerability, cross-system union and a graded completeness score bound what remains recoverable; strict PASS alone is dominated by long series.

8 Limitations

One remediator, who also produced the diagnosis; eight documents; two of them effectively unchanged. Re-anchoring used context-based disambiguation for repeated table values (D2: 69, D3: 23 references) — an error there affects relation coverage for that task only, never evidence recall. The answer layer's noise floor is large: on identical evidence the pilot-protocol comparison moved chunking by +6.4 points, so answer-layer deltas below the chunk column's own movement are not evidence of anything. The retrieval layer is the study's reliable instrument. Round 2 ran once per arm with a single seed; the knowledge subset is defined post hoc by term occurrence in gold/intents (D7's 335 "EDU" runs dilute it — excluding them the targeted Δ is larger).

9 Conclusion

The question behind Part 3 was whether a tagged PDF can be made answerable by its author, or whether answerability is a property of the pipeline that reads it. The answer is that it is a property of the pair, and the two halves are not symmetric. Tagging more thoroughly does not help a consumer that already reads the tag tree; it can even hurt when a standard's preferred encoding outruns the consumer. What the author can do is declare relations the document could not previously state — which note qualifies which claim, which headers govern which cell — keep the text layer honest, and attach, as metadata rather than as page content, the expansions and domain knowledge a reader is assumed to bring. These are modest, mechanical edits, and none of them changes what a sighted reader sees.

Their value is realised only downstream. The structure-aware retriever gained nothing from the edits until the pipeline began writing the declared context beside each piece of evidence; once it did, the answer-stage failures that no document edit had been able to touch began to clear, and injected knowledge proved useful precisely where the document used the term. The structure-blind baseline, which has nothing to carry, stayed where the evidence window put it. The practical guidance in §7 follows from this division of labour: a short list of authoring rules for taggers whose payoff is contingent on consumers that carry structure through, and a matching list for the builders of those consumers.

What remains open is scale and generality — eight documents, one remediator, one family of retrievers — and the residual the study could not move: ambiguity in the source itself. Repeated header–value pairs, composite cells and shared markers are fixed in the document or not at all, and the answerability lint that finds them is, in the end, the most direct tool this line of work offers an author. Part 2 measured the distance between compliant and answerable; Part 3 shows where that distance lies and that a good part of it can be closed from both ends.

Artifacts: dtaa-eval/phase3-arms.md, dtaa-eval/phase3-knowledge.md, dtaa-eval/abbreviations.json, dtaa-eval/phase3-summary.md, dtaa-eval/phase3-retrieval.md, dtaa-runs/corpus-pilot-improved/, dtaa-runs/corpus-perfect-b1400/, dtaa/reanchor.py, dtaa/context.py, docs/dtaa-phase3-plan.md, docs/dtaa-phase3-results.md.