3.2 Model Selection: Haiku, Sonnet, and Opus
3.2.1 A Consistent Ordering: Haiku, Sonnet, Opus
Anthropic's model family spans a spectrum of capability, speed, and cost, and the names map to a consistent ordering you can rely on across the whole exam. Haiku sits at the fast, low-cost end -- built for speed and efficiency on straightforward tasks. Sonnet is the balanced middle -- strong general capability at moderate cost and speed, the everyday workhorse for most business tasks. Opus sits at the far end -- the most capable tier, reserved for the deepest reasoning on genuinely complex problems, at the highest cost.
| Model | Position | Strengths | Typical fit |
|---|---|---|---|
| Haiku | Fastest, lowest cost | Speed and efficiency on straightforward tasks | High-volume, simple, latency-sensitive work (short replies, classification, quick extraction) |
| Sonnet | Balanced | Strong general capability at moderate cost/speed | The everyday workhorse for most business tasks |
| Opus | Most capable, highest cost | Deepest reasoning on complex problems | Hard analysis, nuanced or multi-step reasoning where quality justifies the cost |
A consistent ordering across capability, speed, and cost -- worth memorizing exactly as written.
The one idea to hold onto
More capability generally means higher cost and slower response. There is no single 'best' model in the abstract -- only the best fit for a specific task's needs.
3.2.2 No Universal Best Model -- Only Best Fit
It's tempting to treat Opus as simply "the good one" and default to it whenever there's any doubt. That instinct is exactly backwards. "Always use the most capable model to be safe" wastes cost and latency budget on simple work that a faster, cheaper model would have handled just as well -- there's no safety gained by paying more for capability the task never asked for.
A related trap sits right next to it: assuming a bigger model fixes a bad prompt. A clearer, more specific prompt often matters more than moving up a model tier. Reaching for Opus to compensate for a vague or poorly structured request treats the model tier as a lever for a problem that actually lives in the prompt itself.
3.2.2 -- Exam Trap
'Always use the most capable model to be safe' and 'a bigger model will fix this prompt' are two of the most common exam distractors. Neither holds up: capability beyond what a task needs is pure overspend, and prompt clarity is a separate lever from model tier.
3.2.3 Matching the Model to the Task's Demands
Model selection is about matching a model to what the task actually demands, not to a general sense of importance. Three buckets cover almost every scenario the exam will hand you.
- •Speed/cost matter, reasoning is simple -- a faster, lower-cost model (Haiku). The canonical exam example: generating a high volume of short customer-reply drafts, where speed and cost matter more than deep reasoning. The correct choice is the faster, lower-cost model, not the top-tier one.
- •Balanced everyday work -- Sonnet handles most drafting, summarizing, and analysis well.
- •Hard, high-value reasoning -- Opus, where the quality gain is worth the added cost and time.
3.2.3 -- Key Concept
Match the model to what the task demands: Haiku for simple high-volume/latency-sensitive work, Sonnet for balanced everyday drafting and analysis, Opus for hard, high-value reasoning where the quality gain justifies the cost.
3.2.4 Latency Compounds at Volume
For a single request, using a slower, more capable model just means one slower answer -- a cost that's easy to shrug off. That reasoning breaks down the moment volume enters the picture. For a high-volume batch workflow -- thousands of short replies, for example -- that same per-request slowdown compounds across the whole batch. A top-tier model can bottleneck an entire workflow even though each individual answer it produces is higher quality.
The same per-request slowdown is negligible once and a bottleneck a thousand times over.
3.2.4 -- Exam Trap
Ignoring latency for high-volume tasks -- a slow top-tier model can bottleneck a batch workflow even if each individual answer is higher quality. Also watch for picking a model based on the hardest request a workflow might ever see, rather than the typical request it actually handles.
3.2.5 Thinking in Terms of a Budget
Every task carries an implicit budget across cost, latency, and quality. Selection is the act of matching the model to that target -- not maximizing one dimension (usually quality) while ignoring the other two. A model that is more capable than the task requires is not a safer choice; it is an over-spend of both money and time, with no offsetting benefit for a task that did not need the extra capability.
This reframes the whole domain: model selection isn't about ranking models from worst to best and picking the top of the list. It's about reading the task's actual budget -- how fast does this need to be, how much can it cost, how good does it truly need to be -- and choosing the model that clears that bar without paying for headroom nobody asked for.
3.2.5 -- Key Concept
Every task has an implicit cost/latency/quality budget. Selection matches the model to that target -- an over-capable model for a simple task is an over-spend, not a safer choice.
3.2.6 The Plan/Pricing Tier Boundary
One more constraint sits underneath everything above: which models and features are reachable at all also depends on the plan/pricing tier currently in use. Selection happens within whatever the current plan makes available -- a model that would be the ideal fit for a task may simply be out of reach on a given tier, which is a separate constraint from the cost/latency/quality tradeoff itself.
When cost becomes a real pressure, the correct lever is to right-size the model within the current plan -- not to switch AI platforms or disable product features to cut cost. Those moves address the wrong problem: they change the toolset rather than matching the tool already in hand to the task at hand.
3.2.6 -- Exam Trap
Switching AI platforms or disabling product features to cut cost is the wrong lever -- the real lever available is choosing a right-sized model within the current plan/pricing tier.
Key Takeaways
- ✓Haiku is fastest/cheapest, Sonnet is the balanced everyday workhorse, Opus is most capable and costliest -- a consistent ordering across the family.
- ✓More capability generally means higher cost and slower response; there is no universal 'best' model, only the best fit for the task.
- ✓Defaulting to the top-tier model 'to be safe' wastes cost and latency; a bigger model does not fix a poorly written prompt.
- ✓Simple, high-volume, latency-sensitive work calls for Haiku; balanced everyday work fits Sonnet; hard, high-value reasoning justifies Opus.
- ✓Per-request latency compounds across a high-volume batch workflow -- a slower top-tier model can bottleneck the whole batch.
- ✓Every task has an implicit cost/latency/quality budget; an over-capable model for a simple task is an over-spend, not a safer choice.
- ✓The plan/pricing tier in use bounds which models and features are reachable; right-sizing the model beats switching platforms or disabling features to cut cost.
Check Your Understanding
Test what you learned in this lesson.
Q1.An associate needs a high volume of short customer-reply drafts where speed and cost matter more than deep reasoning. Best choice?
Q2.Which correctly orders the models from lowest cost/fastest to highest capability/cost?
Q3.When is Opus the appropriate choice?
Q4.A workflow generates thousands of short replies per hour. An architect argues for routing every reply through Opus because 'each individual answer will be higher quality.' What is the flaw in this reasoning?
Q5.A team wants to cut their Claude costs and considers switching to a different AI platform or disabling several product features. What does this domain say about that approach?
Practice This Lesson