LLM4MD-LANDSCAPE-BRIEF - Deep Provenance Audit (PUBLIC)

Tier: PUBLIC - published under a per-artifact L7 public-audit exception (publication-package v1.1). Generated: 2026-07-30 by provenance_render.py from the record + session logs. Record-generated, not hand-authored - regenerate rather than edit; edits belong in the record or the session log. Record: the append-only provenance record (LLM4MD-LANDSCAPE-BRIEF) (13 version entries, 3 attributed sessions) Backs: index.html#provenance

How this document was made

Written from scratch in LLM4MD-S02 on a factual base built in S01 — the Nature Medicine paper read in full, its published evaluation pipeline read in full, and five AI-engine research dossiers landed verbatim and adjudicated before any prose was drafted.

It grew from:

SourceRole
Nature Medicine 10.1038/s41591-026-04431-5, full text via PubMed Centralthe subject. Every number in the brief traces here
github.com/nyuolab/clinical-llm-benchmarks, 2,494 lines read in full by subagentthe code audit — three facts the paper does not state; where paper and code disagree the code wins
Perplexity Deep Research dossier — reaction corpusthe public argument: who said what, when, with URLs
Perplexity Computer mode dossier — reaction, 103KBthe richest single input. Vendor responses, adoption figures, the OpenEvidence conflict thread, negative results
Gemini Deep Research dossier — architecturethe Wolters Kluwer FAB platform picture and model-provider evidence
Claude Sonnet dossier — tool statethe product-feature and EHR-integration picture behind the product descriptions
Perplexity Deep Research dossier — practitioner literaturesurrounding literature. Engine output, unverified, largely deferred to Stage 2

Developed across 13 recorded versions (2026-07-26 to 2026-07-30), origin mode fresh.

The public block this backs

Both tiers render from the same record; this is what the reader sees.

Provenance

Owner - Hugh McCutchen

Contributors - Claude Opus 5 — drafting, Hugh McCutchen — direction, Claude Opus 5, Hugh McCutchen — edit passes, Claude Sonnet — blind red team, GPT-5.4 in Perplexity — blind red team, Hugh McCutchen, Hugh McCutchen — caught the source misordering, Hugh McCutchen — called the statistics unusable for the audience, Claude Sonnet — form-comparator scan, Hugh McCutchen — sixteen annotations and the voice decision, Claude Sonnet subagent — chronology validation, Hugh McCutchen — flag dispositions, Hugh McCutchen — caught that the validation record was invisible, Hugh McCutchen — ruled on all three forks

How it began - Written from scratch in LLM4MD-S02 on a factual base built in S01 — the Nature Medicine paper read in full, its published evaluation pipeline read in full, and five AI-engine research dossiers landed verbatim and adjudicated before any prose was drafted. Sourced from: Nature Medicine 10.1038/s41591-026-04431-5, full text via PubMed Central (the subject. Every number in the brief traces here); github.com/nyuolab/clinical-llm-benchmarks, 2,494 lines read in full by subagent (the code audit — three facts the paper does not state; where paper and code disagree the code wins); Perplexity Deep Research dossier — reaction corpus (the public argument: who said what, when, with URLs); Perplexity Computer mode dossier — reaction, 103KB (the richest single input. Vendor responses, adoption figures, the OpenEvidence conflict thread, negative results); Gemini Deep Research dossier — architecture (the Wolters Kluwer FAB platform picture and model-provider evidence); Claude Sonnet dossier — tool state (the product-feature and EHR-integration picture behind the product descriptions); Perplexity Deep Research dossier — practitioner literature (surrounding literature. Engine output, unverified, largely deferred to Stage 2). (origin: fresh)

Type of writing - Inquiry Record

Process - execution, mixed

Tools - Claude Opus 5, my-voice, research-foundation, my-voice banned-patterns, Claude Sonnet, Perplexity, LLM4MD-PROMPT-REDTEAM v1.0, LLM4MD-PROMPT-REDTEAM v2.1, LLM4MD-PROMPT-FORMSCAN v1.0, provenance v2.1, device-write-safety v1.2, session-open-cowork v1.0

Steps run - source work, synthesis, challenge/red-team, verification, customer review, publication fold-back, source verification against the build script

Notes - v0.1: Byte-exact original lost — edited in place. The version chain is broken at v0.1 and is not repairable. | v0.5: The lesson of the arc: a tone problem that survives a sentence-level sweep is an ordering problem. Two blind engines said so independently at the next step. | v0.6-v0.11: The weaker engine could not open any primary and said so only in its closing paragraph; roughly twenty of its UNVERIFIABLE marks were its own retrieval limit. That failure produced the retrieval-declaration requirement in every later prompt. | v0.12-v0.13: Sources 12 and 13 had been inserted above 10 and 11, so the list read 9, 12, 13, 10, 11 and looked as though entries were missing. | v0.14: Two of the round's BLOCKING calls were rejected as engine retrieval failures — PRECISE was verified Completed against the registry record, and the 1,800/1,704 arithmetic reconciles once refusals are counted as items rather than annotations. Both Wolters Kluwer releases were read in full this session to settle the citation. | v0.15: Hugh's objection was to doing statistics a reader cannot check; the counter-argument was that no commentator has made the odds-ratio point, so dropping it would report only one side of a two-sided finding. Resolved by stating the direction without the numbers. | v0.16: The scan opened eight candidate articles and found no published piece combining study explanation, dispute reportage, an adjacent-studies survey and marked opinion. Its structural explanation — that the four moves rest on four different professional warrants which do not stack — is why the piece orients the reader inline rather than leaning on a familiar form. | v0.18: Reportage third-person and impersonal; the marked panels first person plural. The proper noun appears only in the panel label. | v0.20: Body reached its shortest since v0.4 without losing material — the reference moved to an appendix rather than being cut. | v0.21: The 'four months' interval was dropped in both places it appeared: February 2026 is stated without a day anywhere in the evidence base, so the piece was claiming a precision it does not have. The Tytler paragraph was dropped whole rather than neutralised — keeping the observation without his name would absorb a point he made first. | v0.22: The absence claim — no public statement from any of the sixteen authors after 16 June — rests on documented per-author searches in the Computer-mode dossier, and is stated as what the record shows rather than as proof of silence. One concern outlived the flags block and moved into the assembly record: no reviewer has ever opened the figure. | v0.24: The piece argues that its credibility rests on visible method rather than on standing. A validation record compressed to a semicolon contradicted that argument on its own last page. | v0.26: Publication edit list item 1.3 (rewrite the Sources cross-reference §2.1 to §3.1) was NOT applied, deliberately. Verification against the working build showed the edit list was wrong on its own terms: restructure.py performs that sweep at build time and sys.exits if the master has already applied it, and §3.1 does not exist in the master's two-part hierarchy. The same check found that all three MUST-APPLY items were enforced as hard-fail shims in build_llm4md.py (PUB_EDITS), so applying them to the master required removing the shims in the same change — done this session. THE 12 / 13 / 24 DISCREPANCY IS RESOLVED: the three numbers count different things. This record holds 12 VERSION ROWS, several batching a range (v0.6-v0.11, v0.12-v0.13); those rows cover the 24-number chain v0.1-v0.24; handoff v2.1's '13 entries' is the 12 rows plus the seed block. A fourth number nobody had flagged — 'append-only record, 11 entries' in the v0.25 front matter — was also wrong and is corrected. Separately: 22 files sit in ARCHIVE/brief-versions because v0.7 and v0.9 were never archived as their own files; the chain ceiling is still v0.24.

Sources - 3 added across 13 versions (public)

Depth & confidence - working / Deshmukh read in full / unchanged

Level of work - light (13 versions across 3 sessions)

Full history and audit trail available on request.

<!-- rendered from the append-only provenance record (LLM4MD-LANDSCAPE-BRIEF) by provenance_render.py on 2026-07-30 - do not hand-edit -->

Session-by-session record

LLM4MD-S02 (2026-07-26) - v0.1, v0.5, v0.6-v0.11, v0.12-v0.13

Versions produced this session

Session log: the session close log

Decisions locked (mined from the close log)


Deferred / open at close: log is silent.

Unreviewed concerns + judgment calls (mined from the close log)

<!-- END LLM4MD-S02-CLOSE v1.0 -->

LLM4MD-S03 (2026-07-27) - v0.14, v0.15, v0.16, v0.18, v0.20, v0.21, v0.22, v0.24

Versions produced this session

Session log: the session close log

Decisions locked (mined from the close log)

the S02 locked decision. It answers the provenance-class question the first handoff raised — white papers carry the full block by rule — and drops the News mechanics.

panels first person plural; the proper noun only in the panel label.

TAKE, conceding that choosing which caveats to set side by side is itself a view.

the form study's finding: per-claim read-status disclosure is the one thing no comparable published piece does, so publishing the provenance is that discipline carried through.

four-study comparison, the product descriptions, a validated chronology, a glossary, and the six questions that would help.


Deferred / open at close: log is silent.

Unreviewed concerns + judgment calls (mined from the close log)

Its three known defects were found in-house. It is the artifact most likely to be screenshotted and shared without the paper. RES-665 covers the process fix; the figure itself is still unchecked.

the project has and the piece shipped without it.

entirely, v0.18 and v0.19 left byte-identical. Fourth and fifth instances of a class the S02 close already recorded. RES-666.

Research/LLM4MD/ARCHIVE/LLM4MD-S02_TAKE-PANEL_HELD-FOR-LINKEDIN.md. It was step 6 of this session's opening prompt and did not get done.


LLM4MD-S04 (2026-07-30) - v0.26

Versions produced this session

Session log: the session close log

Decisions locked: log is silent.

Deferred / open at close: log is silent.

Unreviewed concerns + judgment calls (mined from the close log)

ConcernHow it resolvedStill needs Hugh?
Asserted "no off-machine copy of any skill" from an empty git remote -vWrong inference — durability is local backup + Drive copy. Hugh corrected in-session. The reasoning error (absence of a remote ≠ absence of durability) is mine and worth watchingNo — but nothing documents Skills durability, so the next session will guess the same way. Offered to record it; not yet done
save_skill called before the disk file was completeViolated commit-then-install. Caught by the closing audit, not by any rule. Logged as an M5 skipped sample rather than glossedNo — RES-705 carries the fix
Edited build_llm4md.py and two skills, all owned by other lanesRES-518 capability-over-territory, Hugh approved the routing up front. Recorded to their owning lanesNo
Left bursera-site-deploy's git-on-mount contradiction unfixedSurgical-edit discipline — v0.4 was scoped to RES-698. Filed as RES-703 insteadYes — RES-703 needs a go
internal/skills/ holds tracked, 20-day-stale skill contentNot deleted — tracked content in another repo is not my callYes — RES-703
Profile v5.0 archived 11:59 today; this session ran on v4.8Same archived-not-applied gap BURSBUILD-S21 caught on v4.9. Second occurrence in four daysYes — now a pattern, not a slip
Left a .probe file at the Work-tree root and in SkillsOpen-gate probe written before the delete grant landed. Both removed once the grant was in placeNo

Corrections & integrity register

Roll-up of every correction, reversal, and integrity event recorded against this artifact. Rows are mined from the record's special_notes and from retrieval-failure / correction markers in the session logs.

Version / sessionWhat happenedWhere it shows in the published work
v0.6-v0.11 / LLM4MD-S02Round-one red team applied. 17 findings accepted, 3 rejected with reason. One BLOCKING factual error corrected. The conclusion moved out of the opening panel into a closing observations section — the structural fix the sentence pass had missed. / The weaker enNot visible — the blocking error and the structural move were corrected before any version left the create lane. First publication was v1.0 on 2026-07-30, from v0.24.
v0.14 / LLM4MD-S03Round-two red team applied in one batch. The paper's odds-ratio result added — it had been in the project's own verified extraction since S01 and had never reached the draft. Wolters Kluwer citation corrected: two of three quoted phrases come from the 7 AprilNot visible — the odds-ratio result was added before first publication. Its absence would have been a material omission had it shipped: the paper's second and harsher measure of the gap.
v0.21 / LLM4MD-S03Final polish. All eleven revision flags closed. The chronology validated by subagent and corrected in five places before it could publish unchecked.Not visible — the five chronology defects were caught by subagent validation pre-publication. The largest was an interval stated as four months the evidence supports only as three to four.
v0.26 / LLM4MD-S04Publication edits folded back from the released v1.0 page so master and site agree. Cut the drafting note addressed to the author ('Word count: 3826... shortest since v0.4'). Corrected 'two working sessions' to three. Reconciled a contradiction inside the ProvMIXED. The two content edits (drafting note cut, 'two'→'three working sessions') were visible on the live page from release, applied by build shim. The master did not carry them until v0.26, so for that window master and site disagreed while the reader saw the correct text. Separately, 'twenty-two versions' vs 'Twenty-four' stood in the published Provenance block until corrected at commit 2e54fdf the same day — reader-visible, and self-contradictory while it stood.

Excluded claims: Three, deliberately. (1) LMIC-cohort observations travel as TECHNIQUE only, never as PREVALENCE — the sources support how a practice is done, not how common it is. (2) Access and concentration — who acquires this skill and what follows if it stays with the powerful — is parked for a separate paper and is not argued here. (3) The surrounding practitioner literature was collected but largely deferred: engine-sourced, unverified against full text, and so excluded from every claim in the piece rather than carried at low confidence. Numbers from that literature are outside the project's standing data lock.

What this does not claim

It does not adjudicate the dispute. Six objections are in play, they are not all compatible with each other, nobody has adjudicated between them and this piece does not either. It does not claim the study is wrong, or that either vendor is right. It does not pick a winner between the two clinical products and the frontier models — the strongest counter-argument, that a single-turn test matches how a physician with ninety seconds actually works, is stated at full strength and left standing. It does not claim a safety difference exists or does not: the study failed to detect one and was not powered to rule one out. It does not assert what either product runs on — the paper's authors say that is not determinable from outside, and every public statement on it, including this piece's, is inference. Depth is uneven by design: deep on the study and its code where every number traces to a primary document, graded on the public argument where four of six commentators reach the piece second-hand, weakest on the four adjacent studies, three of which were not read. The paper's Extended Data was never retrieved and its competing-interests declaration has not been read in the primary.

Recorded depth: working. Confidence: unchanged.


Generated 2026-07-30 from LLM4MD-LANDSCAPE-BRIEF.json + session logs by provenance_render.py. 0 field(s) are silent or need an author - each is marked inline. Close a marker by editing the RECORD, then regenerate; never by editing this page.