In July 2026 I published a LinkedIn post about how I actually work with an LLM. What most readers didn’t see is how the post itself was made: I ran the same conversation in three parallel Claude threads — three models, same prompts, same rules — picked the winning take, and finished the piece with the model that won. The post comes first, exactly as published. The rest of this paper shows the work behind it: the three takes, what each model did under direction, and what each thread cost.
Yesterday I asked an AI “what’s next?” and it answered correctly, because it was reading a plan I wrote. I run research and writing projects inside a chat tool (Claude, in my case). The project holds my plan: the steps, the checkpoints, the working rules for how I want the work done. When I ask what’s next, the AI reads that plan and proposes the next step. Most of the time I approve it. Sometimes I redirect one step, and sometimes the redirect is bigger: that piece of work was wrong from the start.
The comparison I keep coming back to is a head chef. A chef doesn’t cook most of what leaves the kitchen. They design the menu, train the cooks to the house standard, and stand at the pass tasting everything, sending back the plate that is ninety percent right. The food carries their name because the judgment was theirs.
My process has the same shape. I wrote the standards and designed the plan, then gave the AI the coordination job — tracking where we are, keeping the sequence, telling me what’s next — because that is real work and it was never the part I wanted. I kept the design, the taste, and the final say. Every piece of output passes in front of me, and plenty gets sent back.
Most of the energy in AI right now goes to two other approaches. One builds autonomous agents: the machine does the work and the human shrinks to a yes/no at the end, because human time is the cost to cut. The other draws a hard line: ideas and writing stay human, and the AI gets only small mechanical tasks, because ownership is the thing to protect. Both make sense. I’m doing a third thing.
The chef’s seat used to be closed. You didn’t get to run a kitchen without years on the line first. I’m not a trained writer; I’m a curious researcher. This setup lets me produce work I could not have produced alone: researched, checked, versioned, in my own voice, at a standard I set and enforce at every step.
This is one practitioner’s account, not a study. But I’m running the kitchen, and I’m proud of what’s coming out of it.
Think of it as a casting call. One part, three actors reading the same lines, and two of them were fresh off the bus that week — Sonnet 5 and Fable 5 had both just been released. I gave the same brief to three threads in the same Claude project: Sonnet 5 on Medium reasoning effort, Opus 4.8 on Low, and Fable 5 on Low. I throttled everything down on purpose. A LinkedIn post is not complex or wide-ranging work, and I can’t afford to throw maximum horsepower at one.
I honestly hoped Sonnet 5 would win. It’s the everyday workhorse, and a win on the cheap setting would have been the convenient result. That is not what happened. My verdict at the pass, from the session log: “Fable 5 Low wins. I will publish without changing a word.” The most expensive model, on its cheapest setting, produced the winning take. In the end I did keep editing — several more rounds with Fable, pressing it to stop sounding like a bot — and the version above is where that landed. The win stood; the plate still got sent back a few more times before it left the kitchen.
The full three-thread transcript survives, and reading it back, three things stand out.
The bet was overruled at the pass. I expected my usual pick to take it. Instead the plate decided. That matters more than a “cheap model wins” story would have: taste standing at the pass beat my own prior belief about which cook to back, which is a cleaner proof of the chef’s seat than any pricing argument.
The model that went straight to execution was the one I dropped first. Sonnet answered efficiently, then pushed to lock the metaphor and start shaping the deliverable before the idea was settled. I stopped working with it early. The two models I kept stayed in the negotiation — they questioned the framing, pushed back on weak spots, and in one case Opus flagged a risk I had already ruled out, then struck its own pushback when I called it. The failure mode the post describes — the dispatched bot that executes instead of engaging — showed up live, inside the test itself.
The winning draft exists because I said “no — differently.” Mid-exercise I showed each thread a rival’s draft and asked whether anything was worth borrowing. Fable took the best lines on merit and declined the parts that broke my voice rules. Then I sent its own strong draft back once more: it was written for me as the audience, and the readers of the post don’t know my setup. That corrected draft won the pass, and the rounds that followed — the de-botting edits after the transcript ends — produced the published version above. The send-back-the-plate mechanism isn’t described in the transcript. It happens in it.
Two lenses on cost: what the models are priced at, and what the three threads actually consumed.
Official API pricing for the three models as I ran them. Reasoning-effort level does not change the rate — it changes how many tokens a model burns getting to an answer.
| Model · run | Input $/MTok | Output $/MTok | vs. Sonnet |
|---|---|---|---|
| Sonnet 5 · Medium | $2 (intro → $3) | $10 (intro → $15) | 1× |
| Opus 4.8 · Low | $5 | $25 | 2.5× |
| Fable 5 · Low | $10 | $50 | 5× |
Sonnet 5 launched June 30, 2026 at introductory pricing through August 31, 2026. On intro pricing the ratio is 5 / 2.5 / 1; at standard pricing it compresses to roughly 3.3 / 1.7 / 1.
I capture token use with a browser extension that estimates credit consumption; the numbers below are that estimate, not an official meter. One asymmetry to hold while reading: I cut the Sonnet thread off roughly 30% early, so its raw number understates a full run.
| Model · run | Tokens | Credits (est.) | Credits / 1K tokens | Work done |
|---|---|---|---|---|
| Fable 5 · Low | 60,988 | 113,340 | 1,858 | Full run — the winner |
| Opus 4.8 · Low | 74,306 | 71,370 | 960 | Full run, most verbose thread |
| Sonnet 5 · Medium | 46,362 | 24,114 | 520 | Cut ~30% early |
The credits-per-token column is the honest per-model unit cost, independent of how much work each thread did: Sonnet runs about 3.6× cheaper than Fable and about 1.85× cheaper than Opus per unit of work, which matches the shape of the rate card. Adjusting Sonnet upward for the missing 30%, a full Sonnet run still lands under half the Opus thread’s cost, and the winning Fable thread came in at roughly 1.6× Opus.
So the winning take cost about five times the losing one per token, and about 1.6× the runner-up thread in total. For a piece published under my own name, unedited, that trade was worth it. Low effort didn’t dull the judgment; it kept the bill down.
If you are weighing the same three models — Sonnet 5 and Fable 5 brand new, Opus 4.8 the old favorite — this exercise is one honest data point, not a benchmark. For judgment-heavy writing, the model with the best listening won: Fable read my intent back before adding anything, borrowed from a rival draft on merit alone, and rewrote for an audience it was told about once. Effort level mattered less than I expected. Low kept the cost down without dulling the taste.
The finding I trust most isn’t about any one model. Three capable cooks got the same brief, and the difference came from the direction: the plan they were handed, the drafts sent back, the standards already written down before the exercise started. The method carried more of the outcome than the model choice did.
One exercise, one data point.
But the chef tasted every plate, and one came off the pass.