Part 2 measured how much of a document's own task universe a RAG system can complete. Part 3 intervenes on the documents: the eight pilot PDFs were remediated by hand along the Part 2 guidelines (explicit note binding, captions and sectioning, key–value tables as lists, clean text layer), re-compiled into evidence graphs, and re-assessed on the identical frozen tasks and questions. The result separates what document structure can fix from what the retrieval pipeline must fix — and finds that the semantic edits only pay off once the pipeline writes structure into what the model sees.
Part 2 showed that PDFs which pass every accessibility checker are still far from answerable through retrieval-augmented generation. Part 3 asks what to do about it, by intervening on both sides of the gap at once: the eight pilot documents were re-tagged by hand along the Part 2 guidelines — explicit note binding, captions and sectioning, per-cell header pointers, a clean text layer, and an abbreviation list plus a small block of professional knowledge attached as metadata — and the retrieval pipeline was changed so that what a document declares is actually carried to the language model. The frozen tasks, gold and phrasings of Part 2 were reused unchanged, so every difference is attributable to the intervention.
Three observations organise the results. First, structure a retriever already had cannot be improved for it: re-tagging what was already tagged left evidence reach unchanged, and one conformant edit — a caption in place of a heading — briefly made a document worse until the retriever learned to read captions as titles. Second, the levers are declarative: binding notes to the claims they qualify and pointing each data cell at its headers change what the document can state, which shows up as a larger task universe and as unambiguous provenance rather than as a higher score on the old tasks. Third, and decisively, the gain appears only when the pipeline writes the structure next to the evidence: with headers, captions, headings and bound notes serialized beside each value, answer-stage failures fall sharply for the structure-aware system while the structure-blind baseline does not move, and injected domain knowledge helps exactly where its terms occur and nowhere else. From these we propose a short set of authoring guidelines for taggers — header pointers, explicit bindings, additive captions, scoped metadata — that promise to make answerability an attainable property of a tagged PDF, provided consumers carry the structure through.
Part 2 left a specific question open. Eight documents that pass accessibility checkers at 98–100 yielded a task universe of which the best retriever completed 42% of task–phrasing pairs. The paper's §8 turned the diagnosis into eighteen authoring guidelines, but a guideline derived from a failure profile is a hypothesis, not a lever, until a document is changed and the score moves. Part 3 is that test: the same eight PDFs, edited by the same person who tagged them, along the guidelines, with the benchmark side frozen — tasks, gold and the four validated phrasings are reused byte-for-byte.
The failure analysis that preceded the intervention (on the 5,175 pilot runs) also changed what we expected to move. Only 22% of tasks were unanswerable by every system and phrasing; 45 of those 97 had full evidence in front of the model and failed at the answer stage. Half of the retrieval-class failures were context-window losses — the evidence was retrieved but ranked below the shown prefix — not reach failures. And phrasing was the largest single factor (0.55 → 0.26 across the four phrasings of the same task). So the pre-registered prediction (H4b) was that gains, if any, would concentrate in retrieval-class failures on the documents whose punch lists were worked, while the answer stage and lexically distant phrasings would move little.
Remediation was done in the tagging tool under a content-integrity rule: every gold value, list item and note text remains in the document unchanged; headings are not renamed; allowed edits are the guidelines' own — tags, scopes, sectioning, captions, alt text, ActualText, note and link binding. The compiled-graph diff (Table 1) is the authoritative record of what changed.
| document | pp | Table | List Item | Caption | Section | Heading | Note | Figure | figures w/ alt | HEADING_FOR | CAPTION_FOR | REFERENCES |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| D1 FixedIncomeFund fact sheet | 2 | 10 | 10 | 0 | 0→10 | 17 | 0 | 2 | 0→2 | 101→47 | 0 | 0 |
| D2 CorePlusBondFund fact sheet | 4 | 14→13 | 10 | 0→6 | 0→16 | 28→22 | 1 | 6 | 6 | 192→138 | 0→6 | 0 |
| D3 2025 Tax Guide | 5 | 26→21 | 23→44 | 0 | 0 | 38→40 | 34 | 2 | 2 | 250→258 | 0 | 0→58 |
| D4 Ed 43 Summary of Changes | 17 | 36 | 0 | 0 | 0 | 18 | 0 | 1 | 1 | 210 | 0 | 0 |
| D5 Long/Short Growth Equity exposure report | 4 | 17 | 11 | 0 | 0 | 23 | 7 | 3 | 3 | 137 | 0 | 0→7 |
| D6 PIMCO VIT CommodityRealReturn QIR | 13 | 11 | 35 | 0 | 0 | 38 | 7 | 7→6 | 7→6 | 341→339 | 0 | 0→12 |
| D7 Course Outline | 12 | 9 | 54 | 0 | 0 | 45 | 0 | 1 | 1 | 282 | 0 | 0 |
| D8 Hilton FY2024 Slavery & Trafficking Statement | 15 | 7→4 | 44→49 | 1 | 0 | 19 | 0 | 5 | 5 | 195 | 0→1 | 0 |
The edits cluster into four kinds. Explicit note binding (D3, D5, D6): footnotes re-tagged as FENote with reference pointers that the updated dump exports as refer_tag; 58 bindings in the tax guide alone. Captions and sectioning (D1, D2): table blocks wrapped in Sect with a Caption, and — this matters below — in D2 the caption replaced the heading (headings 28→22). Key–value tables as lists (D3, D8): two-column label/value tables converted to L/LI/Lbl/LBody. Text-layer hygiene (D1, D8): kerning artefacts ("F un d") fixed; composite cells given ActualText. D4 and D7 show no tag-tree change and act as controls.
Three of the edits exposed gaps in the consumer, which were fixed so that the intervention is measured against a retriever that understands the structures the guidelines ask for. (a) FENote was unmapped (34 notes in D3 silently became untyped); it now maps to NOTE, and refer_tag compiles to REFERENCES edges (referrer→note, direction-normalised). Expansion follows them both ways, and the NOTE_QUALIFY motif binds by explicit reference first, marker match second. (b) Captions: the compiler now emits CAPTION_FOR (caption→captioned table/figure/list) and, when a container opens with a caption instead of a heading, caption-as-title scope; captions join the structural seed channel and the index's context projection; a caption seed expands into its table. (c) Gold re-anchoring: node ids change on re-tagging, so a re-anchorer maps each gold reference by text (exact → whitespace-squashed → thousands-separator-insensitive → substring → alt text → role/page/ordinal), disambiguated by header/label/heading context, keeping both the original and the new text-layer string as accepted evidence. All 436 tasks re-anchored with no unresolved reference.
Two layers, because they answer different questions. The retrieval layer is deterministic: for every paired run, evidence recall and relation coverage are recomputed from the stored retrieval traces on the full retrieved list (reach) and on the 1,400-word shown prefix. No language model is involved, so the paired delta is exactly the effect of the document edits plus the consumer fixes above. The answer layer repeats Part 2's scoring — answerer over the shown evidence, judge over proposition checklists, strict PASS contract — but the pilot had run with an 8,000-character window (one chunk shown) while the current protocol shows 1,400 words (two chunks). A first pass against the pilot scores produced a +6.4-point "gain" for chunking on byte-identical evidence; the original substrate was therefore re-answered and re-judged at 1,400 words so that before and after share one protocol. That re-baseline is the comparison reported here; the pilot-protocol comparison is kept only as a measure of answerer/judge drift.
Statistics: exact McNemar on discordant pairs per system; a task-clustered bootstrap CI on the mean ΔPASS; per-document McNemar for the document table. Pre-registered hypotheses are in docs/dtaa-phase3-plan.md.
| system | evidence recall (reach) | full evidence (reach) | relation coverage (reach) | evidence recall (shown) | full evidence (shown) |
|---|---|---|---|---|---|
| chunk hybrid | 0.823→0.827 (+0.004) | 0.656→0.667 [0→1 20, 1→0 0, p=1.9e-06] | 0.768→0.780 | 0.608→0.612 | 0.405→0.411 |
| Egres | 0.846→0.846 (+0.000) | 0.699→0.701 [0→1 16, 1→0 13, p=0.71] | 0.760→0.758 | 0.806→0.808 | 0.646→0.650 |
| Egres (reconstructed) | 0.600→0.604 (+0.004) | 0.344→0.354 [0→1 17, 1→0 0, p=1.5e-05] | 0.485→0.493 | 0.539→0.543 | 0.299→0.309 |
| document | chunk hybrid | Egres | Egres (reconstructed) |
|---|---|---|---|
| D1 | 0.94→0.98 (+4) | 0.88→0.89 (+1) | 0.58→0.61 (+4) |
| D2 | 0.85→0.85 (+0) | 0.86→0.87 (+1) | 0.50→0.50 (+0) |
| D3 | 0.76→0.76 (+0) | 0.93→0.93 (-0) | 0.73→0.73 (+0) |
| D4 | 0.96→0.96 (+0) | 0.91→0.91 (+0) | 0.73→0.73 (+0) |
| D5 | 0.91→0.91 (+0) | 0.67→0.67 (+0) | 0.58→0.58 (+0) |
| D6 | 0.82→0.82 (+0) | 0.73→0.73 (+0) | 0.51→0.51 (+0) |
| D7 | 0.67→0.67 (+0) | 0.83→0.83 (+0) | 0.55→0.55 (+0) |
| D8 | 0.79→0.79 (+0) | 0.95→0.93 (-2) | 0.67→0.67 (+0) |
The intervention is neutral at the retrieval layer for the structure-preserving retriever and for the reconstructed substrate; the structure-blind chunker gains where the text layer was cleaned (D1: kerning). The only per-document movement of size was D2, where turning headings into captions removed the tables from their section scope: Egres reach fell nine points on that document until captions were compiled and consumed as titles, after which it is one point up. Every motif-level Δ lies within ±3 points.
| system | runs | At before | At after | Δ pts | 0→1 | 1→0 | McNemar p | 95% CI | evidence recall | words shown |
|---|---|---|---|---|---|---|---|---|---|---|
| chunk hybrid | 1725 | 0.358 | 0.364 | +0.6 | 10 | 0 | 0.002 | [+0.1, +1.3] | 0.61→0.61 | 1162→1162 |
| Egres | 1725 | 0.544 | 0.547 | +0.2 | 28 | 24 | 0.68 | [-0.7, +1.3] | 0.81→0.81 | 1042→1041 |
| Egres (reconstructed) | 1725 | 0.261 | 0.271 | +1.0 | 17 | 0 | 1.5e-05 | [+0.3, +1.9] | 0.54→0.54 | 1131→1131 |
| document | chunk hybrid | Egres | Egres (reconstructed) |
|---|---|---|---|
| D1 FixedIncomeFund fact sheet | 0.55→0.60 (+5*) | 0.64→0.63 (-1) | 0.26→0.35 (+9*) |
| D2 CorePlusBondFund fact sheet | 0.51→0.51 (+0) | 0.63→0.70 (+7*) | 0.29→0.29 (+0) |
| D3 2025 Tax Guide | 0.27→0.27 (+0) | 0.62→0.60 (-2) | 0.38→0.38 (+0) |
| D4 Ed 43 Summary of Changes | 0.39→0.39 (+0) | 0.65→0.65 (+0) | 0.37→0.37 (+0) |
| D5 Long/Short Growth Equity exposure report | 0.27→0.27 (+0) | 0.34→0.34 (+0) | 0.15→0.15 (+0) |
| D6 PIMCO VIT CommodityRealReturn QIR | 0.31→0.31 (+0) | 0.34→0.34 (+1) | 0.13→0.13 (+0) |
| D7 Course Outline | 0.35→0.35 (+0) | 0.61→0.61 (+0) | 0.33→0.33 (+0) |
| D8 Hilton FY2024 Slavery & Trafficking Statement | 0.25→0.25 (+0) | 0.48→0.45 (-3*) | 0.14→0.14 (+0) |
| failure class at baseline | chunk hybrid n / before→after (Δ) | Egres n / before→after (Δ) | Egres (reconstructed) n / before→after (Δ) |
|---|---|---|---|
| R | 1005 / 0.00→0.01 (+1) | 596 / 0.00→0.02 (+2) | 1192 / 0.00→0.01 (+1) |
| A | 77 / 0.00→0.00 (+0) | 138 / 0.00→0.09 (+9) | 61 / 0.00→0.00 (+0) |
| G | 1 / 0.00→0.00 (+0) | 30 / 0.00→0.00 (+0) | 3 / 0.00→0.00 (+0) |
| Q | 24 / 0.00→0.00 (+0) | 22 / 0.00→0.09 (+9) | 18 / 0.00→0.00 (+0) |
| PASS | 618 / 1.00→1.00 (+0) | 939 / 1.00→0.97 (-3) | 451 / 1.00→1.00 (+0) |
| variant | chunk hybrid | Egres | Egres (reconstructed) |
|---|---|---|---|
| Q0 | 433 / 0.42→0.43 (+1) | 433 / 0.64→0.64 (-0) | 433 / 0.30→0.32 (+1) |
| Q1 | 433 / 0.38→0.39 (+1) | 433 / 0.59→0.60 (+1) | 433 / 0.27→0.29 (+1) |
| Q2 | 429 / 0.28→0.28 (+0) | 429 / 0.38→0.38 (-1) | 429 / 0.18→0.18 (+0) |
| Q3 | 430 / 0.35→0.35 (+0) | 430 / 0.57→0.58 (+1) | 430 / 0.29→0.30 (+1) |
| motif | chunk hybrid | Egres | Egres (reconstructed) |
|---|---|---|---|
| FIGURE_ALT | 60 / 0.00→0.00 (+0) | 60 / 0.50→0.50 (+0) | 60 / 0.00→0.00 (+0) |
| HEADING_LIST | 76 / 0.49→0.49 (+0) | 76 / 0.63→0.63 (+0) | 76 / 0.33→0.33 (+0) |
| HEADING_SCOPE | 213 / 0.54→0.54 (+0) | 213 / 0.78→0.77 (-1) | 213 / 0.60→0.60 (+0) |
| LABEL_VALUE | 165 / 0.56→0.56 (+0) | 165 / 0.62→0.59 (-3) | 165 / 0.30→0.30 (+0) |
| NOTE_QUALIFY | 36 / 0.22→0.22 (+0) | 36 / 0.44→0.42 (-3) | 36 / 0.06→0.06 (+0) |
| SECTION_SET | 111 / 0.32→0.32 (+0) | 111 / 0.56→0.59 (+3) | 111 / 0.21→0.21 (+0) |
| TABLE_CELL | 360 / 0.29→0.30 (+1) | 360 / 0.49→0.49 (-0) | 360 / 0.22→0.24 (+2) |
| TABLE_COMPARE | 252 / 0.33→0.36 (+2) | 252 / 0.60→0.62 (+1) | 252 / 0.21→0.23 (+1) |
| TABLE_SERIES | 452 / 0.32→0.32 (+0) | 452 / 0.41→0.43 (+2) | 452 / 0.20→0.21 (+1) |
| system | At pilot (8,000 chars) | At original @1,400 words | Δ budget (pts) | At improved @1,400 words | Δ remediation (pts) |
|---|---|---|---|---|---|
| chunk hybrid | 0.30 | 0.36 | +5.8 | 0.36 | +0.6 |
| Egres | 0.42 | 0.54 | +12.4 | 0.55 | +0.2 |
| Egres (reconstructed) | 0.23 | 0.26 | +3.1 | 0.27 | +1.0 |
Table 4b is the study's least expected result. Re-answering the original documents under the 1,400-word window instead of the pilot's 8,000-character one raised Egres from 0.42 to 0.54 and chunking from 0.30 to 0.36; the remediation then added +0.2 and +0.6 points respectively. Half of Egres' retrieval-class failures in the pilot were evidence that had been retrieved but not shown; widening the window is what converted them.
Reading Tables 4–8 against Table 2: whatever moves at the answer layer moved without the evidence set changing, so the chunk-hybrid column is the floor for answerer/judge variance under this protocol, and the Egres column should be read net of it (H4c: Egres − chunk ≈ -0.3 points). The retrieved-but-not-shown share of Egres' class-R failures is 135/596 before and 123/591 after — the budget, not the document, decides those.
| document | eligible tasks | AMBIGUOUS rejections | NOTE_QUALIFY tasks (explicitly bound) |
|---|---|---|---|
| D1 | 291→293 | 0→0 | 0→0 (0 explicit) |
| D2 | 344→345 | 0→0 | 0→0 (0 explicit) |
| D3 | 270→279 | 0→0 | 2→18 (18 explicit) |
| D4 | 889→889 | 0→0 | 0→0 (0 explicit) |
| D5 | 2,068→2,070 | 2→2 | 4→6 (4 explicit) |
| D6 | 713→714 | 131→131 | 3→3 (3 explicit) |
| D7 | 84→85 | 0→0 | 0→0 (0 explicit) |
| D8 | 52→49 | 10→10 | 0→0 (0 explicit) |
Explicit binding is the one guideline whose effect is visible at the universe level: the tax guide's deterministically bindable qualification tasks go from 2 to 18 because the repeated */*** markers that defeated the page-unique rule are now resolved by the document itself. The ambiguity lint did not move (D6: 131 repeated header–value pairs; D8: 10) — captions and summaries were added to D2, where nothing was flagged, not to the documents where the lint fired; and converting D8's tables to lists removed three LABEL_VALUE tasks without disambiguating the flagged pairs.
Egres reads the tag tree. Adding Sect wrappers, captions and list tags to a document whose headings, tables and lists were already tagged gives it nothing new to walk; reach stays at 0.846. The edits change how the document describes its parts, not how they are connected — and description is consumed at the answer stage and in the index, neither of which was reading it (§5.4).
Replacing a heading with a caption is the correct PDF/UA form for a titled table block, and D2 lost nine points of reach for it. The lesson generalises: a structure standard's preferred encoding is only as good as the least capable consumer, and the guideline should say "add a caption; keep the heading" until consumers treat captions as titles. The consumer fix is small (CAPTION_FOR edges, caption seeds, caption scope) and should be part of any structure-aware retriever.
The only edit that changed what the document can declare is note binding. Its effect is invisible on the frozen tasks — they were sampled from the universe the old document could generate, so they were bindable already — and visible on the universe the new document generates (2→18). The evaluation lesson is that an intervention study must re-enumerate, not only re-score: frozen tasks measure regression, the universe measures capability.
The answerer receives each evidence item as [id] (TABLE_CELL p2) 16.2: no row or column header, no caption, no heading path, no bound note. The wrong-neighbour reads that make up 135 of Egres' 200 answer-stage failures are the direct consequence, and no document edit can reach them. Likewise the index embeds ActualText | text | Alt only, so expansion text (E), table Summary and abbreviation lists — the guidelines' vehicles for phrasing robustness — never affect a seed. The remediation was aimed at the two failure classes the pipeline was structurally unable to pass on to the model.
Round 2 acts on §5.4. On the tagging-tool side, every TD now carries explicit Headers pointers (compiled at top priority, origin SOURCE_ATTRIBUTE: D4 2,154/2,154 header edges declared, D1 387/387 — against 261 declared + 53 positional in the original D1), and an abbreviation list plus a block of reliable professional knowledge (the credit-rating hierarchy AAA … CAA/CCC as used by Moody's, S&P and Fitch) is attached to each document's metadata — no PDF structure involved. On the pipeline side, three arms on the improved substrate share documents, frozen tasks, questions, retrieval traces and an 1,800-word budget and differ only in what is written next to each evidence item: A bare item text; B + structural context from compiled edges (heading path, caption, row/column headers, list label, bound notes); C + abbreviation expansions and knowledge glossed onto the items that contain the term, and embedded into the index projection. The chunk baseline has no edges and receives no glosses, so it is the control for both factors.
At the retrieval layer the declared headers change nothing (Egres reach 0.846 in both) — the compiler was already deriving the same relations from Scope — so every effect below is an answer-layer effect.
| arm | chunk hybrid At / class-A share of failures / Q2 | Egres At / class-A share of failures / Q2 | Egres (reconstructed) At / class-A share of failures / Q2 |
|---|---|---|---|
| orig@1400 | 0.358 / 0.07 / 0.28 | 0.544 / 0.18 / 0.38 | 0.261 / 0.05 / 0.18 |
| impr@1400 | 0.364 / 0.07 / 0.28 | 0.547 / 0.17 / 0.38 | 0.271 / 0.05 / 0.18 |
| A plain@1800 | 0.480 / 0.10 / 0.43 | 0.563 / 0.19 / 0.39 | 0.288 / 0.05 / 0.20 |
| B +context | 0.478 / 0.11 / 0.42 | 0.595 / 0.13 / 0.42 | 0.292 / 0.04 / 0.20 |
| C +context+knowledge | 0.478 / 0.11 / 0.42 | 0.598 / 0.15 / 0.42 | 0.291 / 0.05 / 0.20 |
| factor | chunk hybrid Δ pts (0→1 / 1→0, McNemar p) | Egres Δ pts (0→1 / 1→0, McNemar p) | Egres (reconstructed) Δ pts (0→1 / 1→0, McNemar p) |
|---|---|---|---|
| budget 1,400 → 1,800 words (A − impr@1400) | +11.6 (220 / 20, p=9.1e-44) | +1.6 (35 / 7, p=1.5e-05) | +1.6 (32 / 4, p=1.9e-06) |
| cell/heading/caption/note context (B − A) | -0.2 (18 / 21, p=0.75) | +3.2 (69 / 14, p=6.8e-10) | +0.4 (26 / 19, p=0.37) |
| abbreviations + professional knowledge (C − B) | +0.0 (0 / 0, p=1) | +0.3 (43 / 38, p=0.66) | -0.1 (11 / 13, p=0.84) |
| both (C − A) | -0.2 (18 / 21, p=0.75) | +3.5 (95 / 35, p=1.4e-07) | +0.3 (25 / 19, p=0.45) |
Table-cell context is the lever. Writing headers, captions, heading paths and bound notes next to the value moves Egres by +3.2 points (p = 7e-10) while the chunk baseline moves -0.2: 46% of Egres' answer-stage (class A) failures at arm A pass at arm B, notes +11 and section sets +10 points, the drug-formulary tables of D4 +15, and the lexically distant Q2 phrasing +3. This is the effect the round-1 remediation was aiming at and could not produce, because the structure never reached the model.
| system | tasks | runs | At B → C | Δ pts | 0→1 / 1→0 | McNemar p |
|---|---|---|---|---|---|---|
| chunk hybrid | touching a glossed term | 757 | 0.474 → 0.474 | +0.0 | 0 / 0 | 1 |
| chunk hybrid | not touching one | 968 | 0.481 → 0.481 | +0.0 | 0 / 0 | 1 |
| Egres | touching a glossed term | 757 | 0.596 → 0.620 | +2.4 | 28 / 10 | 0.0051 |
| Egres | not touching one | 968 | 0.594 → 0.581 | -1.3 | 15 / 28 | 0.066 |
| Egres (reconstructed) | touching a glossed term | 757 | 0.318 → 0.310 | -0.8 | 1 / 7 | 0.07 |
| Egres (reconstructed) | not touching one | 968 | 0.272 → 0.276 | +0.4 | 10 / 6 | 0.45 |
| document (terms) | chunk hybrid n / B→C | Egres n / B→C | Egres (reconstructed) n / B→C |
|---|---|---|---|
| D1 | 30 / 1.00→1.00 | 30 / 0.80→0.97 | 30 / 0.40→0.37 |
| D2 | 44 / 0.75→0.75 | 44 / 0.70→0.75 | 44 / 0.45→0.45 |
| D3 | 55 / 0.07→0.07 | 55 / 0.45→0.49 | 55 / 0.05→0.05 |
| D4 | 149 / 0.66→0.66 | 149 / 0.87→0.85 | 149 / 0.49→0.47 |
| D5 | 55 / 0.62→0.62 | 55 / 0.29→0.44 | 55 / 0.04→0.02 |
| D6 | 69 / 0.33→0.33 | 69 / 0.22→0.23 | 69 / 0.17→0.16 |
| D7 | 335 / 0.38→0.38 | 335 / 0.61→0.61 | 335 / 0.36→0.36 |
| D8 | 20 / 0.40→0.40 | 20 / 0.35→0.45 | 20 / 0.00→0.00 |
Professional knowledge works where it applies, and costs a little where it does not. Pooled over all runs the knowledge arm is flat (Egres +0.3). On the 757 Egres runs whose task involves a glossed term it is +2.4 points (p = 0.0051): the credit-quality tasks of D1 go from 0.80 to 0.97 once "BAA" is accompanied by "medium grade, lowest investment-grade tier", D5's exposure tasks from 0.29 to 0.44 (net/gross/market-cap definitions), D8 from 0.35 to 0.45. On the remaining runs it is -1.3 (p = 0.066): gloss lines spend budget. The chunk baseline is exactly 0 in both subsets (identical evidence, cached answers). Knowledge should therefore be attached per item, only where the term occurs — which is what the metadata-driven design does — not appended globally.
The budget, again. Raising the shared window from 1,400 to 1,800 words moves the chunk baseline +11.6 points (a third chunk fits) and Egres +1.6. Under the same budget Egres with context (0.59) leads chunking (0.48) by twelve points on the improved documents.
What the two rounds support, stated as instructions. The evidence behind each item is the table it cites; items without a measured effect are marked as such.
T1 — Give every data cell explicit header pointers. Set Headers on each TD (all row and column headers that apply, including group headers), not only Scope on the TH. Retrieval reach does not change (the compiler derives the same relations from Scope), but the declared pointers are what let a consumer print "row ‘Chile’ · column ‘Fund’" beside a value with certainty, and they resolve repeated values unambiguously (D2: 69 ambiguous gold references → 4). Round 2 rests on them.
T2 — Bind notes explicitly. Tag footnotes as FENote (or Note) with reference pointers from the marker to the note; never rely on a bare * being unique on the page. This is the one edit that expanded what the document can declare (D3: 2 → 18 bindable qualification tasks, Table 9) and it is what lets the consumer show "note: …" beside the claim (NOTE_QUALIFY +11 points in round 2).
T3 — Caption every table and figure, and keep the heading. A Caption as first child of the table block is right; a caption instead of a heading removed the table from its section scope for a heading-only consumer (D2 −9 points of reach, §5.2). Until every consumer treats captions as titles, add — do not replace.
T4 — Make the caption say what the table is. "Country distribution (%), fund vs. index, as of 31 Mar 2025" beats "Table 3". Captions are indexed and shown as context; a caption that repeats the document title adds nothing (Table 1's D2 captions did, and the gain came from the scope, not the words).
T5 — Put units and periods in headers, one value per cell. "$1,192.50 plus 12%" in one cell, or a unit that lives only in a footnote, defeats numeric matching and comparison. Composite cells were among the round-1 answer-stage residuals; splitting them is a source edit no pipeline can replicate.
T6 — Keep the text layer machine-clean. Fix kerning artefacts ("F un d"), keep ActualText consistent with the visible string (the "7,566" vs "7566" case broke string matching for every system), and avoid ActualText that changes the number's form. This was the only document edit that moved the structure-blind baseline (D1 +4 reach).
T7 — Disambiguate repeated header–value pairs. When the same header pair maps to different values in one section (D6: 131 cases, D8: 10), add a caption or a qualifying header ("2Q24", "Class Y"); the answerability lint lists them. No consumer fixed these and none can.
T8 — Attach an abbreviation list as document metadata. Every "BAA", "DIN/PIN", "MAGI", "PVITINST" the document uses, with its expansion — none of the eight files carried an E attribute. Round 2 shows expansions and short domain definitions attached per term are worth +2.4 points on the tasks that use them (Table 12); the same text placed nowhere is worth nothing.
T9 — Attach reliable professional knowledge, scoped to the terms that occur. Standard hierarchies (rating tiers, dosage forms, share-class codes) belong in metadata, not in the page. They helped exactly where the term appears (D1 credit quality 0.80 → 0.97) and cost budget where it does not (−1.3 on unrelated tasks); the consumer must apply them per item.
T10 — Key–value tables as lists only when the pairs are self-contained. Converting two-column label/value tables to L/LI/Lbl/LBody is legitimate, but it removed three tasks in D8 and cost 3 points there (Table 5): a list body loses the column header that told the reader what kind of value it is. If the table had meaningful column headers, keep it a table with T1.
T11 — Section headings that match the reading. An H nested under the wrong parent ("Key employee" under "Highly compensated employee") produces gold the reader disagrees with and answers the judge rejects. Levels should follow the visual hierarchy, and every Sect should open with its own heading or caption.
T12 — Informative alt text, or none. Figure tasks are answerable only through alt text (chunk baseline 0.00 on FIGURE_ALT); "Brandmark of …" answers nothing. State what the figure conveys, with its numbers.
T13 — Do not re-tag what already works. Wrapping already-tagged headings, tables and lists in more containers changed nothing for a structure-aware retriever (§5.1). Spend the effort on T1, T2, T5–T9.
P1 — Serialize structure into the evidence block. Heading path, caption, row/column headers, list label and bound notes next to each item: +3.2 points for Egres, none for chunks (Table 11). The cheapest change in the study and the largest.
P2 — Apply glosses per item, never globally. Expansions and knowledge only where the term occurs; a global preamble spends the window on unrelated tasks (−1.3 where the term is absent).
P3 — Budget per evidence bundle. A cell travels with its headers and caption, a claim with its note; trim bundles, not items. Report "retrieved but not shown" separately: it was half of the retrieval failures, and the 1,400 → 1,800-word step alone was worth +11.6 points to the chunk baseline.
P4 — Read captions as titles and follow reference edges. CAPTION_FOR and REFERENCES are small additions to a heading-only walker and they are what T2/T3 assume.
P5 — Treat declared relations as provenance. Keep the origin of every header edge (declared / scope-derived / positional); it tells the reader how much to trust a pairing.
E1 — Re-enumerate, not only re-score. Frozen tasks measure regression; the universe measures capability (T2's effect was invisible on frozen tasks).
E2 — Hold the protocol fixed and keep a control. Half the first before/after was a budget change; the structure-blind baseline on identical evidence is the noise floor.
E3 — Report At with its companions. Any-phrasing answerability, cross-system union and a graded completeness score bound what remains recoverable; strict PASS alone is dominated by long series.
One remediator, who also produced the diagnosis; eight documents; two of them effectively unchanged. Re-anchoring used context-based disambiguation for repeated table values (D2: 69, D3: 23 references) — an error there affects relation coverage for that task only, never evidence recall. The answer layer's noise floor is large: on identical evidence the pilot-protocol comparison moved chunking by +6.4 points, so answer-layer deltas below the chunk column's own movement are not evidence of anything. The retrieval layer is the study's reliable instrument. Round 2 ran once per arm with a single seed; the knowledge subset is defined post hoc by term occurrence in gold/intents (D7's 335 "EDU" runs dilute it — excluding them the targeted Δ is larger).
The question behind Part 3 was whether a tagged PDF can be made answerable by its author, or whether answerability is a property of the pipeline that reads it. The answer is that it is a property of the pair, and the two halves are not symmetric. Tagging more thoroughly does not help a consumer that already reads the tag tree; it can even hurt when a standard's preferred encoding outruns the consumer. What the author can do is declare relations the document could not previously state — which note qualifies which claim, which headers govern which cell — keep the text layer honest, and attach, as metadata rather than as page content, the expansions and domain knowledge a reader is assumed to bring. These are modest, mechanical edits, and none of them changes what a sighted reader sees.
Their value is realised only downstream. The structure-aware retriever gained nothing from the edits until the pipeline began writing the declared context beside each piece of evidence; once it did, the answer-stage failures that no document edit had been able to touch began to clear, and injected knowledge proved useful precisely where the document used the term. The structure-blind baseline, which has nothing to carry, stayed where the evidence window put it. The practical guidance in §7 follows from this division of labour: a short list of authoring rules for taggers whose payoff is contingent on consumers that carry structure through, and a matching list for the builders of those consumers.
What remains open is scale and generality — eight documents, one remediator, one family of retrievers — and the residual the study could not move: ambiguity in the source itself. Repeated header–value pairs, composite cells and shared markers are fixed in the document or not at all, and the answerability lint that finds them is, in the end, the most direct tool this line of work offers an author. Part 2 measured the distance between compliant and answerable; Part 3 shows where that distance lies and that a good part of it can be closed from both ends.