PrepGenAICerts
Domain 3: Product and Model SelectionLesson 11 of 26

3.2 Model Selection: Haiku, Sonnet, and Opus

3.2.1 A Consistent Ordering: Haiku, Sonnet, Opus

Anthropic's model family spans a spectrum of capability, speed, and cost, and the names map to a consistent ordering you can rely on across the whole exam. Haiku sits at the fast, low-cost end -- built for speed and efficiency on straightforward tasks. Sonnet is the balanced middle -- strong general capability at moderate cost and speed, the everyday workhorse for most business tasks. Opus sits at the far end -- the most capable tier, reserved for the deepest reasoning on genuinely complex problems, at the highest cost.

ModelPositionStrengthsTypical fit
HaikuFastest, lowest costSpeed and efficiency on straightforward tasksHigh-volume, simple, latency-sensitive work (short replies, classification, quick extraction)
SonnetBalancedStrong general capability at moderate cost/speedThe everyday workhorse for most business tasks
OpusMost capable, highest costDeepest reasoning on complex problemsHard analysis, nuanced or multi-step reasoning where quality justifies the cost

A consistent ordering across capability, speed, and cost -- worth memorizing exactly as written.

ℹ️

The one idea to hold onto

More capability generally means higher cost and slower response. There is no single 'best' model in the abstract -- only the best fit for a specific task's needs.

3.2.2 No Universal Best Model -- Only Best Fit

It's tempting to treat Opus as simply "the good one" and default to it whenever there's any doubt. That instinct is exactly backwards. "Always use the most capable model to be safe" wastes cost and latency budget on simple work that a faster, cheaper model would have handled just as well -- there's no safety gained by paying more for capability the task never asked for.

A related trap sits right next to it: assuming a bigger model fixes a bad prompt. A clearer, more specific prompt often matters more than moving up a model tier. Reaching for Opus to compensate for a vague or poorly structured request treats the model tier as a lever for a problem that actually lives in the prompt itself.

⚠️

3.2.2 -- Exam Trap

'Always use the most capable model to be safe' and 'a bigger model will fix this prompt' are two of the most common exam distractors. Neither holds up: capability beyond what a task needs is pure overspend, and prompt clarity is a separate lever from model tier.

3.2.3 Matching the Model to the Task's Demands

Model selection is about matching a model to what the task actually demands, not to a general sense of importance. Three buckets cover almost every scenario the exam will hand you.

  • Speed/cost matter, reasoning is simple -- a faster, lower-cost model (Haiku). The canonical exam example: generating a high volume of short customer-reply drafts, where speed and cost matter more than deep reasoning. The correct choice is the faster, lower-cost model, not the top-tier one.
  • Balanced everyday work -- Sonnet handles most drafting, summarizing, and analysis well.
  • Hard, high-value reasoning -- Opus, where the quality gain is worth the added cost and time.

3.2.3 -- Key Concept

Match the model to what the task demands: Haiku for simple high-volume/latency-sensitive work, Sonnet for balanced everyday drafting and analysis, Opus for hard, high-value reasoning where the quality gain justifies the cost.

3.2.4 Latency Compounds at Volume

For a single request, using a slower, more capable model just means one slower answer -- a cost that's easy to shrug off. That reasoning breaks down the moment volume enters the picture. For a high-volume batch workflow -- thousands of short replies, for example -- that same per-request slowdown compounds across the whole batch. A top-tier model can bottleneck an entire workflow even though each individual answer it produces is higher quality.

One slow answer vs. a slow batchSingle requestone slower answereasy to absorbHigh-volume batchsame slowdown x thousandsbottlenecks the whole workflowa top-tier model's per-request cost compounds at volume even when each answer is better

The same per-request slowdown is negligible once and a bottleneck a thousand times over.

⚠️

3.2.4 -- Exam Trap

Ignoring latency for high-volume tasks -- a slow top-tier model can bottleneck a batch workflow even if each individual answer is higher quality. Also watch for picking a model based on the hardest request a workflow might ever see, rather than the typical request it actually handles.

3.2.5 Thinking in Terms of a Budget

Every task carries an implicit budget across cost, latency, and quality. Selection is the act of matching the model to that target -- not maximizing one dimension (usually quality) while ignoring the other two. A model that is more capable than the task requires is not a safer choice; it is an over-spend of both money and time, with no offsetting benefit for a task that did not need the extra capability.

This reframes the whole domain: model selection isn't about ranking models from worst to best and picking the top of the list. It's about reading the task's actual budget -- how fast does this need to be, how much can it cost, how good does it truly need to be -- and choosing the model that clears that bar without paying for headroom nobody asked for.

3.2.5 -- Key Concept

Every task has an implicit cost/latency/quality budget. Selection matches the model to that target -- an over-capable model for a simple task is an over-spend, not a safer choice.

3.2.6 The Plan/Pricing Tier Boundary

One more constraint sits underneath everything above: which models and features are reachable at all also depends on the plan/pricing tier currently in use. Selection happens within whatever the current plan makes available -- a model that would be the ideal fit for a task may simply be out of reach on a given tier, which is a separate constraint from the cost/latency/quality tradeoff itself.

When cost becomes a real pressure, the correct lever is to right-size the model within the current plan -- not to switch AI platforms or disable product features to cut cost. Those moves address the wrong problem: they change the toolset rather than matching the tool already in hand to the task at hand.

⚠️

3.2.6 -- Exam Trap

Switching AI platforms or disabling product features to cut cost is the wrong lever -- the real lever available is choosing a right-sized model within the current plan/pricing tier.

Key Takeaways

  • Haiku is fastest/cheapest, Sonnet is the balanced everyday workhorse, Opus is most capable and costliest -- a consistent ordering across the family.
  • More capability generally means higher cost and slower response; there is no universal 'best' model, only the best fit for the task.
  • Defaulting to the top-tier model 'to be safe' wastes cost and latency; a bigger model does not fix a poorly written prompt.
  • Simple, high-volume, latency-sensitive work calls for Haiku; balanced everyday work fits Sonnet; hard, high-value reasoning justifies Opus.
  • Per-request latency compounds across a high-volume batch workflow -- a slower top-tier model can bottleneck the whole batch.
  • Every task has an implicit cost/latency/quality budget; an over-capable model for a simple task is an over-spend, not a safer choice.
  • The plan/pricing tier in use bounds which models and features are reachable; right-sizing the model beats switching platforms or disabling features to cut cost.

Check Your Understanding

Test what you learned in this lesson.

Q1.An associate needs a high volume of short customer-reply drafts where speed and cost matter more than deep reasoning. Best choice?

Q2.Which correctly orders the models from lowest cost/fastest to highest capability/cost?

Q3.When is Opus the appropriate choice?

Q4.A workflow generates thousands of short replies per hour. An architect argues for routing every reply through Opus because 'each individual answer will be higher quality.' What is the flaw in this reasoning?

Q5.A team wants to cut their Claude costs and considers switching to a different AI platform or disabling several product features. What does this domain say about that approach?

Practice This Lesson

PrepGenAICerts.com is an independent third-party exam-prep platform for the Claude Certified Architect (CCA-F) certification. We are not affiliated with, endorsed by, or acting on behalf of Anthropic PBC.

Note: New premium upgrades are temporarily paused while we resolve an issue with our payment provider. Existing premium members retain full access.