HumanLedFull Provenance Below

The Chef Post

One brief, three models, one palate — the published piece and the bake-off behind it
Hugh McCutchen · v1.2 · July 2026
Part One

The published post

In July 2026 I published a LinkedIn post about how I actually work with an LLM. What most readers didn’t see is how the post itself was made: I ran the same conversation in three parallel Claude threads — three models, same prompts, same rules — picked the winning take, and finished the piece with the model that won. The post comes first, exactly as published. The rest of this paper shows the work behind it: the three takes, what each model did under direction, and what each thread cost.

Yesterday I asked an AI “what’s next?” and it answered correctly, because it was reading a plan I wrote. I run research and writing projects inside a chat tool (Claude, in my case). The project holds my plan: the steps, the checkpoints, the working rules for how I want the work done. When I ask what’s next, the AI reads that plan and proposes the next step. Most of the time I approve it. Sometimes I redirect one step, and sometimes the redirect is bigger: that piece of work was wrong from the start.

The comparison I keep coming back to is a head chef. A chef doesn’t cook most of what leaves the kitchen. They design the menu, train the cooks to the house standard, and stand at the pass tasting everything, sending back the plate that is ninety percent right. The food carries their name because the judgment was theirs.

My process has the same shape. I wrote the standards and designed the plan, then gave the AI the coordination job — tracking where we are, keeping the sequence, telling me what’s next — because that is real work and it was never the part I wanted. I kept the design, the taste, and the final say. Every piece of output passes in front of me, and plenty gets sent back.

Most of the energy in AI right now goes to two other approaches. One builds autonomous agents: the machine does the work and the human shrinks to a yes/no at the end, because human time is the cost to cut. The other draws a hard line: ideas and writing stay human, and the AI gets only small mechanical tasks, because ownership is the thing to protect. Both make sense. I’m doing a third thing.

The chef’s seat used to be closed. You didn’t get to run a kitchen without years on the line first. I’m not a trained writer; I’m a curious researcher. This setup lets me produce work I could not have produced alone: researched, checked, versioned, in my own voice, at a standard I set and enforce at every step.

This is one practitioner’s account, not a study. But I’m running the kitchen, and I’m proud of what’s coming out of it.

The writer as head chefDiagram of the writer's role as head chef: the chef keeps menu design, taste, and veto; the kitchen holds the clipboard and the brigade; work and questions cross the pass in both directions. The chef stays in the kitchen What the writer keeps, what the kitchen runs, and the pass between them The writer — head chef Keeps what cannot be delegated: menu design · taste · veto Wrote every standard the kitchen runs on the pass — every plate crosses here, in both directions the plan · the standards "no — re-fire that, differently" "the next step is…" "what do you need from me, chef?" The kitchen — trained to the house The clipboard sequencing · tracking answering "what's next" The brigade research · drafting checking its own work The dish carries the chef's name because the judgment was theirs · July 2026
THE WRITER AS HEAD CHEF — WHAT STAYS WITH THE WRITER, WHAT THE KITCHEN RUNS, AND THE PASS BETWEEN THEM
Part Two

The bake-off — how the post got made

Think of it as a casting call. One part, three actors reading the same lines, and two of them were fresh off the bus that week — Sonnet 5 and Fable 5 had both just been released. I gave the same brief to three threads in the same Claude project: Sonnet 5 on Medium reasoning effort, Opus 4.8 on Low, and Fable 5 on Low. I throttled everything down on purpose. A LinkedIn post is not complex or wide-ranging work, and I can’t afford to throw maximum horsepower at one.

I honestly hoped Sonnet 5 would win. It’s the everyday workhorse, and a win on the cheap setting would have been the convenient result. That is not what happened. My verdict at the pass, from the session log: “Fable 5 Low wins. I will publish without changing a word.” The most expensive model, on its cheapest setting, produced the winning take. In the end I did keep editing — several more rounds with Fable, pressing it to stop sounding like a bot — and the version above is where that landed. The win stood; the plate still got sent back a few more times before it left the kitchen.

Three models, one kitchenComparison of three parallel Claude threads used to draft one LinkedIn post, with the same human judgment applied across all three. White background for export. Three models, one kitchen, one published post Same prompts run in three parallel threads — the chef tasted every plate Sonnet 5 · Medium "Want me to push on which of these is the strongest vehicle for a post?" What happened · Fast, execution-first · Drove at the deliverable · Dropped after two turns Not the job for this seat Opus 4.8 · Low "It wasn't a gap; it was your sentence. So strike that pushback." What happened · Hardest critic of the frame · Restated my point as a risk · Owned the miss when caught Best sparring partner Fable 5 · Low — published "The empty seat isn't 'human in the loop' — it's design, taste, and veto." What happened · Read back before adding · Took the rival's best lines · Rewrote it for cold readers Published, word for word The same chef worked all three kitchens "What's next?" hand over the clipboard "Nope — not that, differently." the veto at the pass "What was my last sentence?" send the critic's plate back Same prompts, three parallel Claude threads · one human palate · July 2026
THREE MODELS, ONE KITCHEN — SAME BRIEF, THREE PARALLEL THREADS, ONE HUMAN PALATE AT THE PASS

What the transcript shows

The full three-thread transcript survives, and reading it back, three things stand out.

The bet was overruled at the pass. I expected my usual pick to take it. Instead the plate decided. That matters more than a “cheap model wins” story would have: taste standing at the pass beat my own prior belief about which cook to back, which is a cleaner proof of the chef’s seat than any pricing argument.

The model that went straight to execution was the one I dropped first. Sonnet answered efficiently, then pushed to lock the metaphor and start shaping the deliverable before the idea was settled. I stopped working with it early. The two models I kept stayed in the negotiation — they questioned the framing, pushed back on weak spots, and in one case Opus flagged a risk I had already ruled out, then struck its own pushback when I called it. The failure mode the post describes — the dispatched bot that executes instead of engaging — showed up live, inside the test itself.

The winning draft exists because I said “no — differently.” Mid-exercise I showed each thread a rival’s draft and asked whether anything was worth borrowing. Fable took the best lines on merit and declined the parts that broke my voice rules. Then I sent its own strong draft back once more: it was written for me as the audience, and the readers of the post don’t know my setup. That corrected draft won the pass, and the rounds that followed — the de-botting edits after the transcript ends — produced the published version above. The send-back-the-plate mechanism isn’t described in the transcript. It happens in it.

A footnote on the method: the post about running a kitchen was made by running a kitchen. Same brief, three kitchens, one pass, one palate. I tasted all three, sent two back, and kept the third at the pass until it was right.
Part Three

What it cost

Two lenses on cost: what the models are priced at, and what the three threads actually consumed.

The rate card

Official API pricing for the three models as I ran them. Reasoning-effort level does not change the rate — it changes how many tokens a model burns getting to an answer.

Model · runInput $/MTokOutput $/MTokvs. Sonnet
Sonnet 5 · Medium$2 (intro → $3)$10 (intro → $15)
Opus 4.8 · Low$5$252.5×
Fable 5 · Low$10$50

Sonnet 5 launched June 30, 2026 at introductory pricing through August 31, 2026. On intro pricing the ratio is 5 / 2.5 / 1; at standard pricing it compresses to roughly 3.3 / 1.7 / 1.

What the three threads actually consumed

I capture token use with a browser extension that estimates credit consumption; the numbers below are that estimate, not an official meter. One asymmetry to hold while reading: I cut the Sonnet thread off roughly 30% early, so its raw number understates a full run.

Model · runTokensCredits (est.)Credits / 1K tokensWork done
Fable 5 · Low60,988113,3401,858Full run — the winner
Opus 4.8 · Low74,30671,370960Full run, most verbose thread
Sonnet 5 · Medium46,36224,114520Cut ~30% early

The credits-per-token column is the honest per-model unit cost, independent of how much work each thread did: Sonnet runs about 3.6× cheaper than Fable and about 1.85× cheaper than Opus per unit of work, which matches the shape of the rate card. Adjusting Sonnet upward for the missing 30%, a full Sonnet run still lands under half the Opus thread’s cost, and the winning Fable thread came in at roughly 1.6× Opus.

So the winning take cost about five times the losing one per token, and about 1.6× the runner-up thread in total. For a piece published under my own name, unedited, that trade was worth it. Low effort didn’t dull the judgment; it kept the bill down.

Part Four

One case, one data point

If you are weighing the same three models — Sonnet 5 and Fable 5 brand new, Opus 4.8 the old favorite — this exercise is one honest data point, not a benchmark. For judgment-heavy writing, the model with the best listening won: Fable read my intent back before adding anything, borrowed from a rival draft on merit alone, and rewrote for an audience it was told about once. Effort level mattered less than I expected. Low kept the cost down without dulling the taste.

The finding I trust most isn’t about any one model. Three capable cooks got the same brief, and the difference came from the direction: the plan they were handed, the drafts sent back, the standards already written down before the exercise started. The method carried more of the outcome than the model choice did.

One exercise, one data point.

But the chef tasted every plate, and one came off the pass.

Provenance

ModelSource session: Claude Sonnet 5 · Medium, Claude Opus 4.8 · Low, Claude Fable 5 · Low (three parallel threads, 2026-07-02; winning take by Fable 5 · Low, finished over further edit rounds with Fable before publication) · Web edition: Claude Fable 5 (claude-fable-5) · Medium (2026-07-07) · Placement (SVG figures, og:image, hub-card image): Claude Sonnet 5 (claude-sonnet-5) · BURSBUILD-S09 (2026-07-08)
Versionv1.2 · HTML · July 2026
SessionBURSFORGE-S01 · 2026-07-07 · source bake-off run 2026-07-02 · placed BURSBUILD-S09 · 2026-07-08
AudienceLeaders evaluating how to direct LLMs for serious writing; readers choosing between Sonnet 5, Opus 4.8, and Fable 5 and weighing effort levels
SourcesGoogle Doc “Linked-in Chef story - 3 models help” — the complete three-thread transcript (~16,800 words, read in full per RES-314), containing the published post, all drafts, Hugh’s session verdict, the official rate-card lookup, and the browser-extension cost capture; Linear RES-346 (build ticket), RES-353 (placement); white-paper structural exemplar claude-guide/WPP-S11-WHITEPAPER-v2.2.html; two SVG diagrams + two JPEG crops (og:image, hub-card) supplied and directed by Hugh at placement (BURSBUILD-S09)
Decision / ActionAuthored as ONE piece (published post + bake-off + cost analysis) per the hub cluster description; RES-346’s story/technical split option flagged, not exercised. Published post carried verbatim. At placement: two SVG diagrams inserted; og:image + hub-card JPEG crops added in a second pass, framed per Hugh's direction (Fable green card only for og:image; title + writer/chef gold box for the hub card), bordered in var(--gray-rule).
Iteration notesv1.2: section-nav added (RES-313 opt-in); table row-hover removed, solid row backgrounds + explicit ink. v1.1: post body replaced with the published final. v1.0 (web): first HTML edition. Placement (BURSBUILD-S09, same v1.2): two SVG figures inserted; og:image + hub-card image added in a follow-up edit, body text otherwise byte-faithful to the BURSFORGE submission throughout.
AssumptionsCredit figures are browser-extension estimates, labeled as such; official API rates as published 2026-07 (Sonnet 5 intro pricing sunsets 2026-08-31). Plan-usage weights are Anthropic’s qualitative descriptions, not exact multipliers.
Scope exclusionsThe full draft-by-draft transcript (summarized, not reproduced); Opus 4.8’s consolation post; general model benchmarking beyond this one exercise.
Tool chainGoogle Drive read (source transcript); file tools + shell placement (RES-315 drill); my-voice v3.1 (mode 2, banned-patterns sweep); provenance skill; PIL (image crop/border/resize)
Review statusv1.2 — placed BURSBUILD-S09; SVG figures + og:image + hub-card image all confirmed with Hugh in-session; awaiting deploy.bat