Field Notes
Share:
Copy link
Copied!
Why the same team runs four different Claude models
I sent a stack of expense reports through Opus last week, out of habit more than anything, and the bill came back oversized for what the work actually was. The task was categorizing line items by GL code and flagging anything missing a receipt. That's categorization work, not reconciliation work, and I was paying reconciliation rates for it.
It's the mistake you make when you treat "which model" as a single dial running from cheap to expensive, rather than four distinct tools built for four distinct kinds of work. Once you see the actual axis, the choice mostly makes itself.
Cost tracks the mess, not the stakes
The question that actually matters isn't "how important is this task." It's "how much conflicting, unstructured material has to get reconciled before a judgment call is even possible." A single clean dataset with an obvious ask stays cheap regardless of how much money is riding on the answer. A pile of inconsistent source documents that need cross-referencing is where a bigger model earns its cost — because that reconciliation work is exactly what separates the tiers.
Anthropic currently ships four tiers on that axis: Haiku 4.5, Sonnet 5, Opus 4.8, and Fable 5, each priced roughly 2–3x the one before it. The spectrum below lays them out in order, with what each one actually costs per million tokens processed. (There's a fifth model, Mythos 5, that shares Fable's weights but is locked to approved cybersecurity and biosecurity partners through Project Glasswing — not something you'll be routing a close package to.)

What the tiers cost, in real numbers
The bar above shows the sticker price, and the ratio inside it tells you something worth remembering: output tokens cost five times input tokens at every single tier. That matters more than people account for — a model that reasons out loud or drafts long memos burns through output tokens fast, so the "cheap" tier on a verbose task can end up costing more than the "expensive" tier on a terse one. The chart below breaks input and output pricing apart by tier so the gap is easier to see at a glance.

Context window tells a second part of the story the chart doesn't show. Haiku caps out at 200,000 tokens; Sonnet, Opus, and Fable all carry a full 1 million. That's the real reason Haiku stays confined to triage work — it's not just cheaper, it's structurally unfit for a data room. You can't reconcile forty contracts against a cap table in a window that can't hold forty contracts and a cap table at once.
What each tier actually does in a finance seat
Haiku's job is volume without judgment. Take a batch of vendor invoices with messy, inconsistent line-item descriptions — some abbreviated, some duplicated, none using the same category names twice. Feed Haiku the batch and have it normalize every line into a consistent set of GL categories, flag anything that doesn't map cleanly, and hand back a clean table. No reasoning happens here worth paying for, and the volume is exactly what Haiku's built for.
Sonnet is the daily analyst. Give it the actuals-vs-budget export for the monthly close, last month's commentary for tone, and a note on what drove the big swings — headcount timing, a one-time legal fee. It drafts variance commentary in your house style and, this is the part that matters, tells you which variances it couldn't explain from the data alone rather than papering over the gap. You still own whether a variance is a trend or noise. The blank page just gets smaller.
Opus is where the stakes and the mess both go up. A buy-side diligence workstream — forty contracts, a cap table, three years of financials, a management deck — asking Opus to reconcile customer concentration figures across all of it and flag anywhere the numbers don't tie, is the associate who rereads the contract looking for what everyone else missed. That behavior shows up more reliably here than at Sonnet's tier, specifically because the task spans documents that actually disagree with each other.
Fable is the outer edge, and it earns that edge on scope, not on any single hard question. An IC memo build that ingests the full data room, researches the competitor landscape, pulls prior memos for house style, and produces a complete sourced first draft with an explicit list of open diligence questions — that's a genuinely long-horizon, high-volume-input task where Opus can lose the thread across hundreds of documents and Sonnet definitely will.
To make the choice concrete, here's the decision I actually run through before opening a chat.

Picking a tier without overthinking it
It's not really about task category — variance commentary, pricing memo, DD teardown — it's about what's sitting in the input. Structured, single-source material with no real ambiguity goes to Haiku. Clean data that needs a competent first draft goes to Sonnet, which is where the bulk of daily finance work should live. Multiple documents with real inconsistencies to catch earns Opus. A genuinely long-horizon synthesis job, the full data room ingested end to end, is the one case worth paying Fable's rate for.
Worth saying plainly: three of these four tiers will handle a routine finance task just fine. The differences show up at the edges, not in the middle of the distribution — which is exactly why defaulting to the top tier out of habit is such an easy trap. Most of what crosses a Strategic Finance desk in a given week is Sonnet work wearing an Opus-sized worry.
I re-ran those expense reports through Haiku the next morning. Same output, a fraction of the cost, and about the same amount of time it took me to notice I'd been doing it wrong.
Published: