LLM4MD-LANDSCAPE-BRIEF - Deep Provenance Audit (PUBLIC)
Tier: PUBLIC - published under a per-artifact L7 public-audit exception (publication-package v1.1). Generated: 2026-07-30 by provenance_render.py from the record + session logs. Record-generated, not hand-authored - regenerate rather than edit; edits belong in the record or the session log. Record: the append-only provenance record (LLM4MD-LANDSCAPE-BRIEF) (13 version entries, 3 attributed sessions) Backs: index.html#provenance
How this document was made
Written from scratch in LLM4MD-S02 on a factual base built in S01 — the Nature Medicine paper read in full, its published evaluation pipeline read in full, and five AI-engine research dossiers landed verbatim and adjudicated before any prose was drafted.
It grew from:
| Source | Role |
|---|---|
| Nature Medicine 10.1038/s41591-026-04431-5, full text via PubMed Central | the subject. Every number in the brief traces here |
| github.com/nyuolab/clinical-llm-benchmarks, 2,494 lines read in full by subagent | the code audit — three facts the paper does not state; where paper and code disagree the code wins |
| Perplexity Deep Research dossier — reaction corpus | the public argument: who said what, when, with URLs |
| Perplexity Computer mode dossier — reaction, 103KB | the richest single input. Vendor responses, adoption figures, the OpenEvidence conflict thread, negative results |
| Gemini Deep Research dossier — architecture | the Wolters Kluwer FAB platform picture and model-provider evidence |
| Claude Sonnet dossier — tool state | the product-feature and EHR-integration picture behind the product descriptions |
| Perplexity Deep Research dossier — practitioner literature | surrounding literature. Engine output, unverified, largely deferred to Stage 2 |
Developed across 13 recorded versions (2026-07-26 to 2026-07-30), origin mode fresh.
The public block this backs
Both tiers render from the same record; this is what the reader sees.
Provenance
Owner - Hugh McCutchen
Contributors - Claude Opus 5 — drafting, Hugh McCutchen — direction, Claude Opus 5, Hugh McCutchen — edit passes, Claude Sonnet — blind red team, GPT-5.4 in Perplexity — blind red team, Hugh McCutchen, Hugh McCutchen — caught the source misordering, Hugh McCutchen — called the statistics unusable for the audience, Claude Sonnet — form-comparator scan, Hugh McCutchen — sixteen annotations and the voice decision, Claude Sonnet subagent — chronology validation, Hugh McCutchen — flag dispositions, Hugh McCutchen — caught that the validation record was invisible, Hugh McCutchen — ruled on all three forks
How it began - Written from scratch in LLM4MD-S02 on a factual base built in S01 — the Nature Medicine paper read in full, its published evaluation pipeline read in full, and five AI-engine research dossiers landed verbatim and adjudicated before any prose was drafted. Sourced from: Nature Medicine 10.1038/s41591-026-04431-5, full text via PubMed Central (the subject. Every number in the brief traces here); github.com/nyuolab/clinical-llm-benchmarks, 2,494 lines read in full by subagent (the code audit — three facts the paper does not state; where paper and code disagree the code wins); Perplexity Deep Research dossier — reaction corpus (the public argument: who said what, when, with URLs); Perplexity Computer mode dossier — reaction, 103KB (the richest single input. Vendor responses, adoption figures, the OpenEvidence conflict thread, negative results); Gemini Deep Research dossier — architecture (the Wolters Kluwer FAB platform picture and model-provider evidence); Claude Sonnet dossier — tool state (the product-feature and EHR-integration picture behind the product descriptions); Perplexity Deep Research dossier — practitioner literature (surrounding literature. Engine output, unverified, largely deferred to Stage 2). (origin: fresh)
Type of writing - Inquiry Record
Process - execution, mixed
Tools - Claude Opus 5, my-voice, research-foundation, my-voice banned-patterns, Claude Sonnet, Perplexity, LLM4MD-PROMPT-REDTEAM v1.0, LLM4MD-PROMPT-REDTEAM v2.1, LLM4MD-PROMPT-FORMSCAN v1.0, provenance v2.1, device-write-safety v1.2, session-open-cowork v1.0
Steps run - source work, synthesis, challenge/red-team, verification, customer review, publication fold-back, source verification against the build script
Notes - v0.1: Byte-exact original lost — edited in place. The version chain is broken at v0.1 and is not repairable. | v0.5: The lesson of the arc: a tone problem that survives a sentence-level sweep is an ordering problem. Two blind engines said so independently at the next step. | v0.6-v0.11: The weaker engine could not open any primary and said so only in its closing paragraph; roughly twenty of its UNVERIFIABLE marks were its own retrieval limit. That failure produced the retrieval-declaration requirement in every later prompt. | v0.12-v0.13: Sources 12 and 13 had been inserted above 10 and 11, so the list read 9, 12, 13, 10, 11 and looked as though entries were missing. | v0.14: Two of the round's BLOCKING calls were rejected as engine retrieval failures — PRECISE was verified Completed against the registry record, and the 1,800/1,704 arithmetic reconciles once refusals are counted as items rather than annotations. Both Wolters Kluwer releases were read in full this session to settle the citation. | v0.15: Hugh's objection was to doing statistics a reader cannot check; the counter-argument was that no commentator has made the odds-ratio point, so dropping it would report only one side of a two-sided finding. Resolved by stating the direction without the numbers. | v0.16: The scan opened eight candidate articles and found no published piece combining study explanation, dispute reportage, an adjacent-studies survey and marked opinion. Its structural explanation — that the four moves rest on four different professional warrants which do not stack — is why the piece orients the reader inline rather than leaning on a familiar form. | v0.18: Reportage third-person and impersonal; the marked panels first person plural. The proper noun appears only in the panel label. | v0.20: Body reached its shortest since v0.4 without losing material — the reference moved to an appendix rather than being cut. | v0.21: The 'four months' interval was dropped in both places it appeared: February 2026 is stated without a day anywhere in the evidence base, so the piece was claiming a precision it does not have. The Tytler paragraph was dropped whole rather than neutralised — keeping the observation without his name would absorb a point he made first. | v0.22: The absence claim — no public statement from any of the sixteen authors after 16 June — rests on documented per-author searches in the Computer-mode dossier, and is stated as what the record shows rather than as proof of silence. One concern outlived the flags block and moved into the assembly record: no reviewer has ever opened the figure. | v0.24: The piece argues that its credibility rests on visible method rather than on standing. A validation record compressed to a semicolon contradicted that argument on its own last page. | v0.26: Publication edit list item 1.3 (rewrite the Sources cross-reference §2.1 to §3.1) was NOT applied, deliberately. Verification against the working build showed the edit list was wrong on its own terms: restructure.py performs that sweep at build time and sys.exits if the master has already applied it, and §3.1 does not exist in the master's two-part hierarchy. The same check found that all three MUST-APPLY items were enforced as hard-fail shims in build_llm4md.py (PUB_EDITS), so applying them to the master required removing the shims in the same change — done this session. THE 12 / 13 / 24 DISCREPANCY IS RESOLVED: the three numbers count different things. This record holds 12 VERSION ROWS, several batching a range (v0.6-v0.11, v0.12-v0.13); those rows cover the 24-number chain v0.1-v0.24; handoff v2.1's '13 entries' is the 12 rows plus the seed block. A fourth number nobody had flagged — 'append-only record, 11 entries' in the v0.25 front matter — was also wrong and is corrected. Separately: 22 files sit in ARCHIVE/brief-versions because v0.7 and v0.9 were never archived as their own files; the chain ceiling is still v0.24.
Sources - 3 added across 13 versions (public)
Depth & confidence - working / Deshmukh read in full / unchanged
Level of work - light (13 versions across 3 sessions)
Full history and audit trail available on request.
<!-- rendered from the append-only provenance record (LLM4MD-LANDSCAPE-BRIEF) by provenance_render.py on 2026-07-30 - do not hand-edit -->
Session-by-session record
LLM4MD-S02 (2026-07-26) - v0.1, v0.5, v0.6-v0.11, v0.12-v0.13
Versions produced this session
- v0.1 - First draft, written against the S01 factual base. _Contributors:_ Claude Opus 5 — drafting, Hugh McCutchen — direction _Notes:_ Byte-exact original lost — edited in place. The version chain is broken at v0.1 and is not repairable.
- v0.5 - Sentence-level tone pass — nine constructed contrasts and every body-prose em-dash removed. Reported as fixing the register problem; it did not. _Contributors:_ Claude Opus 5, Hugh McCutchen — edit passes _Notes:_ The lesson of the arc: a tone problem that survives a sentence-level sweep is an ordering problem. Two blind engines said so independently at the next step.
- v0.6-v0.11 - Round-one red team applied. 17 findings accepted, 3 rejected with reason. One BLOCKING factual error corrected. The conclusion moved out of the opening panel into a closing observations section — the structural fix the sentence pass had missed. _Contributors:_ Claude Opus 5, Claude Sonnet — blind red team, GPT-5.4 in Perplexity — blind red team, Hugh McCutchen _Notes:_ The weaker engine could not open any primary and said so only in its closing paragraph; roughly twenty of its UNVERIFIABLE marks were its own retrieval limit. That failure produced the retrieval-declaration requirement in every later prompt.
- v0.12-v0.13 - Sources merged into one numeric group with read-status on every entry. Diagram re-axised to v0.7. _Contributors:_ Claude Opus 5, Hugh McCutchen — caught the source misordering _Notes:_ Sources 12 and 13 had been inserted above 10 and 11, so the list read 9, 12, 13, 10, 11 and looked as though entries were missing.
Session log: the session close log
Decisions locked (mined from the close log)
- Publication routes through BURSBUILD. LLM4MD is a create lane. It authors copy; placement and deploy belong to the Bursera Website project via
bursera-site-deploy, and copy crosses as a handoff file into_incoming/. There are no push credentials in a Cowork session in any case. - The piece is a News item, not a white paper. The provenance-class question — News is not in the site's class list, while the BEBURS charter wants a seal on everything published — is BURSBUILD's to answer, not this project's.
- Register: background, not argument. One marked opinion panel at the top, one observations section at the close, reportage in between. Hugh's edits enforced this repeatedly and it is now the piece's governing constraint.
- The "How this was assembled" appendix is internal. Strip before external handoff. Marked in the front matter.
- Two paragraphs are held for the LinkedIn lead, not lost — the PubMed anecdote and the closing context-and-skill paragraph, preserved verbatim in
ARCHIVE/. latest.mdwas not repointed. Standing instruction; the open gate will FAIL on it and that FAIL is expected.
Deferred / open at close: log is silent.
Unreviewed concerns + judgment calls (mined from the close log)
- the governing document §10 is stale the same way the instructions were — it still carries the retired two-claim Part 1 structure and guardrails 2 and 3. It is the governing doc, so this is proposed rather than done.
- The v0.1 brief's byte-exact original is lost. Edited in place; the manifest records 15,495 bytes against 13,442 on disk. Section 3 was preserved from session context, not recovered from disk. Version chain broken at v0.1.
- v0.7 and v0.9 are not separate files. Hugh authored both in place over the v0.6 and v0.8 files, so the filename and the front-matter version disagree for those two.
open_verify.pyfails on a deliberately un-repointed pointer and cannot express the intended state. Not yet ticketed.- The piece is ~4,700 words, well past the 2–3 page Stage-1 target.
<!-- END LLM4MD-S02-CLOSE v1.0 -->
LLM4MD-S03 (2026-07-27) - v0.14, v0.15, v0.16, v0.18, v0.20, v0.21, v0.22, v0.24
Versions produced this session
- v0.14 - Round-two red team applied in one batch. The paper's odds-ratio result added — it had been in the project's own verified extraction since S01 and had never reached the draft. Wolters Kluwer citation corrected: two of three quoted phrases come from the 7 April FAB release, not the 3 June release the sentence cited. Talby quotation restored to full. Code facts scoped to the frontier arm. Three registers defined and marked. _Contributors:_ Claude Opus 5, Claude Sonnet — blind red team, GPT-5.4 in Perplexity — blind red team, Hugh McCutchen _Notes:_ Two of the round's BLOCKING calls were rejected as engine retrieval failures — PRECISE was verified Completed against the registry record, and the 1,800/1,704 arithmetic reconciles once refusals are counted as items rather than annotations. Both Wolters Kluwer releases were read in full this session to settle the citation.
- v0.15 - Panel delimiters paired so a panel can hold a table without silent truncation. Odds-ratio statistics replaced by three plain sentences reporting that a second measure exists and which way it points. 'The authors' reserved for the study's authors throughout. _Contributors:_ Claude Opus 5, Hugh McCutchen — called the statistics unusable for the audience _Notes:_ Hugh's objection was to doing statistics a reader cannot check; the counter-argument was that no commentator has made the odds-ratio point, so dropping it would report only one side of a two-sided finding. Resolved by stating the direction without the numbers.
- v0.16 - Form-comparator findings applied. Orientation moved ahead of the first panel to state position and method. Farah Deshmukh added — she corroborates the conflict allegation independently of the aggregation route and reached the operator-skill point first. _Contributors:_ Claude Opus 5, Claude Sonnet — form-comparator scan, Hugh McCutchen _Notes:_ The scan opened eight candidate articles and found no published piece combining study explanation, dispute reportage, an adjacent-studies survey and marked opinion. Its structural explanation — that the four moves rest on four different professional warrants which do not stack — is why the piece orients the reader inline rather than leaning on a familiar form.
- v0.18 - Voice moved from Hugh personally to Bursera. Sixteen of Hugh's annotations applied. Two sections cut and one closing table reduced; body down 14%. _Contributors:_ Hugh McCutchen — sixteen annotations and the voice decision, Claude Opus 5 _Notes:_ Reportage third-person and impersonal; the marked panels first person plural. The proper noun appears only in the panel label.
- v0.20 - Three registers collapsed to two — the attributed-explanation panel relabelled as the firm's view, conceding that choosing which caveats to set side by side is itself a view. Publishing appendix added: the four-study comparison, the product descriptions, a dated chronology, a glossary, and the six questions that would help. _Contributors:_ Hugh McCutchen, Claude Opus 5 _Notes:_ Body reached its shortest since v0.4 without losing material — the reference moved to an appendix rather than being cut.
- v0.21 - Final polish. All eleven revision flags closed. The chronology validated by subagent and corrected in five places before it could publish unchecked. _Contributors:_ Claude Opus 5, Claude Sonnet subagent — chronology validation, Hugh McCutchen — flag dispositions _Notes:_ The 'four months' interval was dropped in both places it appeared: February 2026 is stated without a day anywhere in the evidence base, so the piece was claiming a precision it does not have. The Tytler paragraph was dropped whole rather than neutralised — keeping the observation without his name would absorb a point he made first.
- v0.22 - A section added for the study's authors, who had been represented only inside the vendors' section. Their standing position and the documented absence of any later statement now sit in their own §2.3. Revision-flags block retired. _Contributors:_ Hugh McCutchen, Claude Opus 5 _Notes:_ The absence claim — no public statement from any of the sixteen authors after 16 June — rests on documented per-author searches in the Computer-mode dossier, and is stated as what the record shows rather than as proof of silence. One concern outlived the flags block and moved into the assembly record: no reviewer has ever opened the figure.
- v0.24 - The rendered provenance block rebuilt around validation. Eight passes now listed individually with who ran each and what it found, plus what was not checked. Previously this was one fragment in a list and one dense paragraph. _Contributors:_ Hugh McCutchen — caught that the validation record was invisible, Claude Opus 5 _Notes:_ The piece argues that its credibility rests on visible method rather than on standing. A validation record compressed to a semicolon contradicted that argument on its own last page.
Session log: the session close log
Decisions locked (mined from the close log)
- Reclassified from News item to white paper on the Research page (Hugh, 2026-07-28). Overrides
the S02 locked decision. It answers the provenance-class question the first handoff raised — white papers carry the full block by rule — and drops the News mechanics.
- The voice is Bursera, not Hugh personally. Reportage third-person and impersonal; the marked
panels first person plural; the proper noun only in the panel label.
- Three registers collapsed to two. The attributed-explanation panel was relabelled BURSERA'S
TAKE, conceding that choosing which caveats to set side by side is itself a view.
- The rendered provenance block publishes. A change from every prior version, and the reason is
the form study's finding: per-claim read-status disclosure is the one thing no comparable published piece does, so publishing the provenance is that discipline carried through.
- A publishing appendix, carrying material cut from the body rather than deleted: the
four-study comparison, the product descriptions, a validated chronology, a glossary, and the six questions that would help.
- Superseded versions are retired by move and rename, never deletion.
Deferred / open at close: log is silent.
Unreviewed concerns + judgment calls (mined from the close log)
- The figure has never been reviewed outside Bursera. Three red-team rounds ran on text alone.
Its three known defects were found in-house. It is the artifact most likely to be screenshotted and shared without the paper. RES-665 covers the process fix; the figure itself is still unchecked.
- Round-three red-team prompt was built and never run. v3.0 is the strongest review instrument
the project has and the piece shipped without it.
- The version chain was broken three more times this session by in-place overwrites — v0.16 lost
entirely, v0.18 and v0.19 left byte-identical. Fourth and fifth instances of a class the S02 close already recorded. RES-666.
- Charter §10 remains stale. Proposed in S02, still not approved.
- The LinkedIn lead was not drafted. Material is held verbatim at
Research/LLM4MD/ARCHIVE/LLM4MD-S02_TAKE-PANEL_HELD-FOR-LINKEDIN.md. It was step 6 of this session's opening prompt and did not get done.
- Stage 2 has not started. Eleven of fifteen items in the research plan's queue are untouched.
LLM4MD-S04 (2026-07-30) - v0.26
Versions produced this session
- v0.26 - Publication edits folded back from the released v1.0 page so master and site agree. Cut the drafting note addressed to the author ('Word count: 3826... shortest since v0.4'). Corrected 'two working sessions' to three. Reconciled a contradiction inside the Provenance block that the publication edit list had not caught: Contributors said 'twenty-two versions' while Level of work said 'Twenty-four' — both now twenty-four. Cover copy (title, tagline, method claim, meta) authored into the master front matter for the first time, plus the share slots. Added a MASTER-TO-BUILD CONTRACT block to the front matter. _Contributors:_ Hugh McCutchen — ruled on all three forks, Claude Opus 5 _Notes:_ Publication edit list item 1.3 (rewrite the Sources cross-reference §2.1 to §3.1) was NOT applied, deliberately. Verification against the working build showed the edit list was wrong on its own terms: restructure.py performs that sweep at build time and sys.exits if the master has already applied it, and §3.1 does not exist in the master's two-part hierarchy. The same check found that all three MUST-APPLY items were enforced as hard-fail shims in build_llm4md.py (PUB_EDITS), so applying them to the master required removing the shims in the same change — done this session. THE 12 / 13 / 24 DISCREPANCY IS RESOLVED: the three numbers count different things. This record holds 12 VERSION ROWS, several batching a range (v0.6-v0.11, v0.12-v0.13); those rows cover the 24-number chain v0.1-v0.24; handoff v2.1's '13 entries' is the 12 rows plus the seed block. A fourth number nobody had flagged — 'append-only record, 11 entries' in the v0.25 front matter — was also wrong and is corrected. Separately: 22 files sit in ARCHIVE/brief-versions because v0.7 and v0.9 were never archived as their own files; the chain ceiling is still v0.24.
Session log: the session close log
Decisions locked: log is silent.
Deferred / open at close: log is silent.
Unreviewed concerns + judgment calls (mined from the close log)
| Concern | How it resolved | Still needs Hugh? |
|---|---|---|
Asserted "no off-machine copy of any skill" from an empty git remote -v | Wrong inference — durability is local backup + Drive copy. Hugh corrected in-session. The reasoning error (absence of a remote ≠ absence of durability) is mine and worth watching | No — but nothing documents Skills durability, so the next session will guess the same way. Offered to record it; not yet done |
save_skill called before the disk file was complete | Violated commit-then-install. Caught by the closing audit, not by any rule. Logged as an M5 skipped sample rather than glossed | No — RES-705 carries the fix |
Edited build_llm4md.py and two skills, all owned by other lanes | RES-518 capability-over-territory, Hugh approved the routing up front. Recorded to their owning lanes | No |
Left bursera-site-deploy's git-on-mount contradiction unfixed | Surgical-edit discipline — v0.4 was scoped to RES-698. Filed as RES-703 instead | Yes — RES-703 needs a go |
internal/skills/ holds tracked, 20-day-stale skill content | Not deleted — tracked content in another repo is not my call | Yes — RES-703 |
| Profile v5.0 archived 11:59 today; this session ran on v4.8 | Same archived-not-applied gap BURSBUILD-S21 caught on v4.9. Second occurrence in four days | Yes — now a pattern, not a slip |
Left a .probe file at the Work-tree root and in Skills | Open-gate probe written before the delete grant landed. Both removed once the grant was in place | No |
Corrections & integrity register
Roll-up of every correction, reversal, and integrity event recorded against this artifact. Rows are mined from the record's special_notes and from retrieval-failure / correction markers in the session logs.
| Version / session | What happened | Where it shows in the published work |
|---|---|---|
| v0.6-v0.11 / LLM4MD-S02 | Round-one red team applied. 17 findings accepted, 3 rejected with reason. One BLOCKING factual error corrected. The conclusion moved out of the opening panel into a closing observations section — the structural fix the sentence pass had missed. / The weaker en | Not visible — the blocking error and the structural move were corrected before any version left the create lane. First publication was v1.0 on 2026-07-30, from v0.24. |
| v0.14 / LLM4MD-S03 | Round-two red team applied in one batch. The paper's odds-ratio result added — it had been in the project's own verified extraction since S01 and had never reached the draft. Wolters Kluwer citation corrected: two of three quoted phrases come from the 7 April | Not visible — the odds-ratio result was added before first publication. Its absence would have been a material omission had it shipped: the paper's second and harsher measure of the gap. |
| v0.21 / LLM4MD-S03 | Final polish. All eleven revision flags closed. The chronology validated by subagent and corrected in five places before it could publish unchecked. | Not visible — the five chronology defects were caught by subagent validation pre-publication. The largest was an interval stated as four months the evidence supports only as three to four. |
| v0.26 / LLM4MD-S04 | Publication edits folded back from the released v1.0 page so master and site agree. Cut the drafting note addressed to the author ('Word count: 3826... shortest since v0.4'). Corrected 'two working sessions' to three. Reconciled a contradiction inside the Prov | MIXED. The two content edits (drafting note cut, 'two'→'three working sessions') were visible on the live page from release, applied by build shim. The master did not carry them until v0.26, so for that window master and site disagreed while the reader saw the correct text. Separately, 'twenty-two versions' vs 'Twenty-four' stood in the published Provenance block until corrected at commit 2e54fdf the same day — reader-visible, and self-contradictory while it stood. |
Excluded claims: Three, deliberately. (1) LMIC-cohort observations travel as TECHNIQUE only, never as PREVALENCE — the sources support how a practice is done, not how common it is. (2) Access and concentration — who acquires this skill and what follows if it stays with the powerful — is parked for a separate paper and is not argued here. (3) The surrounding practitioner literature was collected but largely deferred: engine-sourced, unverified against full text, and so excluded from every claim in the piece rather than carried at low confidence. Numbers from that literature are outside the project's standing data lock.
What this does not claim
It does not adjudicate the dispute. Six objections are in play, they are not all compatible with each other, nobody has adjudicated between them and this piece does not either. It does not claim the study is wrong, or that either vendor is right. It does not pick a winner between the two clinical products and the frontier models — the strongest counter-argument, that a single-turn test matches how a physician with ninety seconds actually works, is stated at full strength and left standing. It does not claim a safety difference exists or does not: the study failed to detect one and was not powered to rule one out. It does not assert what either product runs on — the paper's authors say that is not determinable from outside, and every public statement on it, including this piece's, is inference. Depth is uneven by design: deep on the study and its code where every number traces to a primary document, graded on the public argument where four of six commentators reach the piece second-hand, weakest on the four adjacent studies, three of which were not read. The paper's Extended Data was never retrieved and its competing-interests declaration has not been read in the primary.
Recorded depth: working. Confidence: unchanged.
Generated 2026-07-30 from LLM4MD-LANDSCAPE-BRIEF.json + session logs by provenance_render.py. 0 field(s) are silent or need an author - each is marked inline. Close a marker by editing the RECORD, then regenerate; never by editing this page.