Nature Medicine vs OpenEvidence and UpToDate — Full Provenance and Audit Trail

Document: Nature Medicine Compared the Big Chat Models to OpenEvidence and UpToDate — an inquiry record
Record: LLM4MD-LANDSCAPE-BRIEF — append-only, 12 entries
Published: July 2026

This page is the audit trail behind the provenance block at the foot of the paper. It is rendered from a per-version record kept alongside the drafting and maintained by its author as append-only — entries are added, never rewritten. Nothing on this page is asserted that the record does not carry. Where a red-team finding was rejected rather than accepted, that is recorded too.


How this document was made — in one paragraph

Written from scratch in LLM4MD-S02 on a factual base built in S01 — the Nature Medicine paper read in full, its published evaluation pipeline read in full, and five AI-engine research dossiers landed verbatim and adjudicated before any prose was drafted. 12 entries follow, spanning v0.1 to v0.24. Of the eight validation passes the paper lists, four were run by reviewers outside the drafting: two blind red-team rounds on separate engines that never saw each other's returns, a form study, and a chronology check. Two of the red team's own blocking calls were rejected with reason, and the reasons are recorded. The figure was never reviewed outside Bursera, and the paper says so.


Layer 1 — the summary block this page backs

FieldValue
Document idLLM4MD-LANDSCAPE-BRIEF
OwnerHugh McCutchen
Provenance startfresh — the record opens at the first draft; no earlier work is folded in
Type of writingInquiry Record
Entries recorded12, spanning v0.1 to v0.24. Some entries cover a range of versions where several were cut in one pass
How it beganWritten from scratch in LLM4MD-S02 on a factual base built in S01 — the Nature Medicine paper read in full, its published evaluation pipeline read in full, and five AI-engine research dossiers landed verbatim and adjudicated before any prose was drafted.
  • Nature Medicine 10.1038/s41591-026-04431-5, full text via PubMed Central — the subject. Every number in the brief traces here
  • github.com/nyuolab/clinical-llm-benchmarks, 2,494 lines read in full by subagent — the code audit — three facts the paper does not state; where paper and code disagree the code wins
  • Perplexity Deep Research dossier — reaction corpus — the public argument: who said what, when, with URLs
  • Perplexity Computer mode dossier — reaction, 103KB — the richest single input. Vendor responses, adoption figures, the OpenEvidence conflict thread, negative results
  • Gemini Deep Research dossier — architecture — the Wolters Kluwer FAB platform picture and model-provider evidence
  • Claude Sonnet dossier — tool state — the product-feature and EHR-integration picture behind the product descriptions
  • Perplexity Deep Research dossier — practitioner literature — surrounding literature. Engine output, unverified, largely deferred to Stage 2
Depth, confidence, level of workStated at document level in the paper’s own provenance block, which this page backs rather than restates. The per-entry values below describe each increment, not the whole.

Version-by-version record

0 · v0.1 — 2026-07-26

First draft, written against the S01 factual base.

Contributors: Claude Opus 5 — drafting
Hugh McCutchen — direction
Steps run: source work
synthesis
Notes: Byte-exact original lost — edited in place. The version chain is broken at v0.1 and is not repairable.
This entry: process type: execution; depth: working; confidence in findings: strong on the study, thin on the reaction; level of work: substantial

1 · v0.5 — 2026-07-26

Sentence-level tone pass — nine constructed contrasts and every body-prose em-dash removed. Reported as fixing the register problem; it did not.

Contributors: Claude Opus 5
Hugh McCutchen — edit passes
Steps run: synthesis
Notes: The lesson of the arc: a tone problem that survives a sentence-level sweep is an ordering problem. Two blind engines said so independently at the next step.
This entry: process type: execution; depth: working; confidence in findings: unchanged; level of work: light

2 · v0.6-v0.11 — 2026-07-26

Round-one red team applied. 17 findings accepted, 3 rejected with reason. One BLOCKING factual error corrected. The conclusion moved out of the opening panel into a closing observations section — the structural fix the sentence pass had missed.

Contributors: Claude Opus 5
Claude Sonnet — blind red team
GPT-5.4 in Perplexity — blind red team
Hugh McCutchen
Steps run: source work
synthesis
challenge/red-team
Notes: The weaker engine could not open any primary and said so only in its closing paragraph; roughly twenty of its UNVERIFIABLE marks were its own retrieval limit. That failure produced the retrieval-declaration requirement in every later prompt.
This entry: process type: mixed; depth: deep; confidence in findings: strong on the study and the code, graded on the reaction; level of work: extensive

3 · v0.12-v0.13 — 2026-07-27

Sources merged into one numeric group with read-status on every entry. Diagram re-axised to v0.7.

Contributors: Claude Opus 5
Hugh McCutchen — caught the source misordering
Steps run: synthesis
Notes: Sources 12 and 13 had been inserted above 10 and 11, so the list read 9, 12, 13, 10, 11 and looked as though entries were missing.
This entry: process type: execution; depth: working; confidence in findings: unchanged; level of work: light

4 · v0.14 — 2026-07-27

Round-two red team applied in one batch. The paper's odds-ratio result added — it had been in the project's own verified extraction since S01 and had never reached the draft. Wolters Kluwer citation corrected: two of three quoted phrases come from the 7 April FAB release, not the 3 June release the sentence cited. Talby quotation restored to full. Code facts scoped to the frontier arm. Three registers defined and marked.

Contributors: Claude Opus 5
Claude Sonnet — blind red team
GPT-5.4 in Perplexity — blind red team
Hugh McCutchen
Steps run: source work
synthesis
challenge/red-team
verification
Notes: Two of the round's BLOCKING calls were rejected as engine retrieval failures — PRECISE was verified Completed against the registry record, and the 1,800/1,704 arithmetic reconciles once refusals are counted as items rather than annotations. Both Wolters Kluwer releases were read in full this session to settle the citation.
This entry: process type: mixed; depth: deep; confidence in findings: strong where verified, explicitly graded where not; level of work: extensive

5 · v0.15 — 2026-07-27

Panel delimiters paired so a panel can hold a table without silent truncation. Odds-ratio statistics replaced by three plain sentences reporting that a second measure exists and which way it points. 'The authors' reserved for the study's authors throughout.

Contributors: Claude Opus 5
Hugh McCutchen — called the statistics unusable for the audience
Steps run: synthesis
Notes: Hugh's objection was to doing statistics a reader cannot check; the counter-argument was that no commentator has made the odds-ratio point, so dropping it would report only one side of a two-sided finding. Resolved by stating the direction without the numbers.
This entry: process type: mixed; depth: working; confidence in findings: unchanged; level of work: substantial

6 · v0.16 — 2026-07-27

Form-comparator findings applied. Orientation moved ahead of the first panel to state position and method. Farah Deshmukh added — she corroborates the conflict allegation independently of the aggregation route and reached the operator-skill point first.

Contributors: Claude Opus 5
Claude Sonnet — form-comparator scan
Hugh McCutchen
Steps run: source work
synthesis
challenge/red-team
Notes: The scan opened eight candidate articles and found no published piece combining study explanation, dispute reportage, an adjacent-studies survey and marked opinion. Its structural explanation — that the four moves rest on four different professional warrants which do not stack — is why the piece orients the reader inline rather than leaning on a familiar form.
This entry: process type: mixed; depth: deep; confidence in findings: unchanged; level of work: substantial

7 · v0.18 — 2026-07-27

Voice moved from Hugh personally to Bursera. Sixteen of Hugh's annotations applied. Two sections cut and one closing table reduced; body down 14%.

Contributors: Hugh McCutchen — sixteen annotations and the voice decision
Claude Opus 5
Steps run: synthesis
customer review
Notes: Reportage third-person and impersonal; the marked panels first person plural. The proper noun appears only in the panel label.
This entry: process type: execution; depth: working; confidence in findings: unchanged; level of work: substantial

8 · v0.20 — 2026-07-27

Three registers collapsed to two — the attributed-explanation panel relabelled as the firm's view, conceding that choosing which caveats to set side by side is itself a view. Publishing appendix added: the four-study comparison, the product descriptions, a dated chronology, a glossary, and the six questions that would help.

Contributors: Hugh McCutchen
Claude Opus 5
Steps run: synthesis
customer review
Notes: Body reached its shortest since v0.4 without losing material — the reference moved to an appendix rather than being cut.
This entry: process type: mixed; depth: working; confidence in findings: unchanged; level of work: substantial

9 · v0.21 — 2026-07-27

Final polish. All eleven revision flags closed. The chronology validated by subagent and corrected in five places before it could publish unchecked.

Contributors: Claude Opus 5
Claude Sonnet subagent — chronology validation
Hugh McCutchen — flag dispositions
Steps run: synthesis
challenge/red-team
verification
Notes: The 'four months' interval was dropped in both places it appeared: February 2026 is stated without a day anywhere in the evidence base, so the piece was claiming a precision it does not have. The Tytler paragraph was dropped whole rather than neutralised — keeping the observation without his name would absorb a point he made first.
This entry: process type: execution; depth: working; confidence in findings: strong; citation integrity checked, all 21 sources used and defined; level of work: substantial

10 · v0.22 — 2026-07-28

A section added for the study's authors, who had been represented only inside the vendors' section. Their standing position and the documented absence of any later statement now sit in their own §2.3. Revision-flags block retired.

Contributors: Hugh McCutchen
Claude Opus 5
Steps run: synthesis
customer review
Notes: The absence claim — no public statement from any of the sixteen authors after 16 June — rests on documented per-author searches in the Computer-mode dossier, and is stated as what the record shows rather than as proof of silence. One concern outlived the flags block and moved into the assembly record: no reviewer has ever opened the figure.
This entry: process type: execution; depth: working; confidence in findings: strong where verified; the figure remains unreviewed; level of work: light

11 · v0.24 — 2026-07-28

The rendered provenance block rebuilt around validation. Eight passes now listed individually with who ran each and what it found, plus what was not checked. Previously this was one fragment in a list and one dense paragraph.

Contributors: Hugh McCutchen — caught that the validation record was invisible
Claude Opus 5
Steps run: synthesis
customer review
Notes: The piece argues that its credibility rests on visible method rather than on standing. A validation record compressed to a semicolon contradicted that argument on its own last page.
This entry: process type: execution; depth: working; confidence in findings: unchanged; level of work: light


What was not checked

Carried here from the paper's own provenance block, unchanged. The figure has never been reviewed outside Bursera — all three red-team passes ran on the text alone, and its three known defects were found in-house. The paper's Extended Data was never retrieved. Its competing-interests declaration has not been read in the primary. Three of the four other studies in the appendix were not read. Four commentators' posts were retrieved by an automated pass and never opened at source. Each of those is labelled where it appears in the paper.

A research artifact is only as good as the record behind it. This page is that record. If something here does not reconcile with the paper, the discrepancy is the finding — tell us.

← Return to the paper