PrepGenAICerts
Domain 2: Output Evaluation and ValidationLesson 5 of 26

2.1 Evaluating Accuracy and Completeness

2.1.1 Fluent Is Not the Same as Correct

Claude produces fluent, confident text. That single fact is the reason this entire domain exists, and it's worth sitting with before anything else: the model's writing quality tells you nothing about whether its content is true. A hallucinated statistic reads exactly as smoothly as a verified one. A dropped requirement can be buried inside three paragraphs of otherwise well-organized prose. Before an output is used for anything that matters, it has to be judged against what the task actually required — not against a vague, in-the-moment impression of whether it reads well.

This is the defining mindset of the Associate role, and it's worth naming directly: diligence. The person using Claude stays accountable for the output. Confidence in Claude's wording is never itself evidence that the content is correct. Everything in this domain — checking accuracy, catching hallucinations, fact-checking, deciding when a human needs to sign off, editing, and choosing the right format — is diligence applied at a different stage of the same workflow.

Two ways to evaluate an outputRead draft, then judge (wrong)"does this look okay?"standard shifts to fit the draftunverifiable, inconsistentSet criteria, then judge (right)"what must 'good' satisfy?"decided before reading the draftcheckable, repeatable

Success criteria set before you read the draft keep the standard fixed; letting the draft set the standard is how polish gets mistaken for correctness.

ℹ️

The one idea to hold onto

The core mindset of Domain 2: the person using Claude stays accountable for the output. Confidence in Claude's wording is never itself evidence that the content is correct.

2.1.2 Defining Success Criteria Before You Read the Draft

Anthropic's guidance on defining success criteria applies directly to evaluation: decide up front what "good" looks like, then check the output against that definition. A vague standard like "does this look okay?" can't be verified consistently — two reviewers might disagree, or the same reviewer might miss something on a second read. An explicit, checkable criterion removes the ambiguity. "Must cite the correct regulation subsection." "Must stay under 200 words." "Must not recommend a refund." Each of these can be answered with a yes or no.

The sequencing matters as much as the content of the criteria. Success criteria have to be set before the draft is read, not after. If you wait until an appealing, well-written draft is already in front of you, the draft quietly shapes the standard instead of the standard shaping your judgment of the draft — you'll rationalize why the parts it got right are the parts that mattered.

2.1.2 — Key Concept

Decide what "good" means before you generate or read the output. Explicit, checkable success criteria (word limits, required citations, prohibited content) replace vague "does this look okay?" judgment — and a confident, fluent answer is not evidence that it meets them.

2.1.3 The Five-Part Checklist

Once success criteria are set, a practical five-part checklist turns them into a repeatable review. Run every output-in-question through all five dimensions — not just the one that happens to catch your eye first.

DimensionQuestion it answers
AccuracyAre the facts, figures, names, and citations correct and verifiable?
CompletenessDid it address every part of the request, or quietly drop one?
RelevanceDoes it answer the actual question, not a nearby one?
ConsistencyDo the numbers, claims, and recommendations agree with each other and with the source material?
Fitness for audienceIs the tone, depth, and format right for who will read it?

Each dimension can fail independently — accurate facts don't guarantee completeness, relevance, consistency, or audience fit.

Notice that accuracy is only one-fifth of the checklist. An output can get every individual fact right and still fail the review — by skipping part of a multi-part ask, by answering a slightly different question than the one posed, by contradicting itself between paragraphs, or by being pitched at the wrong reader. The checklist exists precisely because these are independent failure modes, not one problem wearing different masks.

2.1.4 Completeness vs. Sounding Complete

Completeness is the checklist item most often faked by good writing. A response can read as thorough — well-organized, confident, detailed — while quietly skipping one part of a multi-part request. Checking completeness means going back to the original request and confirming each requested part is actually present in the output, not just that the response *feels* comprehensive on a read-through. If someone asked for a summary, a risk assessment, and a recommendation, and the output delivers a beautifully written summary and recommendation but no risk assessment, it is incomplete — no matter how polished the two delivered pieces are.

Relevance has the same trap in a different shape. Claude can produce a fluent, on-topic-sounding answer to a slightly different question than the one that was actually asked. Checking relevance means re-reading the actual question and confirming the output answers *that* question, not a plausible neighbor of it — a common failure when a prompt contains two related sub-questions and the response thoroughly answers only the easier one.

  • Reread the original request line by line and tick off each discrete part — don't rely on a general impression of thoroughness.
  • Restate the actual question in one sentence and check the output answers that sentence, not an adjacent one.
  • Treat a well-organized, confident answer as a reason to look closer, not a reason to skip the check.
⚠️

2.1.4 — Exam Trap

Common exam trap: checking only that the output sounds complete instead of confirming each requested part is present against the original request. A confident, well-written answer is not proof of correctness or completeness.

2.1.5 Discernment: Naming the Skill Behind This Task Statement

Everything covered so far in this lesson -- setting success criteria before reading a draft, running the five-part checklist, catching completeness dressed up as thoroughness, catching relevance dressed up as a nearby answer -- is a single named competency in Anthropic's AI Fluency framework: Discernment, the discipline of judging output against what was actually requested, what the sources actually say, and what the field considers acceptable. The success criteria and the checklist aren't separate tricks bolted onto this lesson; they're Discernment made concrete and checkable.

This course names Diligence explicitly elsewhere -- iteration in Lesson 1.3, ethical use in Lesson 6.4 -- as the competency of staying responsible for an outcome and using AI effectively, ethically, and safely. Discernment is Diligence's quieter partner, and the exam expects you to tell them apart cleanly rather than treat both as "being careful with Claude."

DiscernmentDiligence
Answers the questionHow do I review this well?Why, and when, must I review at all -- and who owns the result?
Shows up asThe five-part checklist, checkable success criteria, tracing a claim to its sourceThe decision to actually run the review, escalate to a human, or disclose AI use, even under time pressure
A failure looks likeMissing a dropped requirement because the checklist wasn't actually appliedSkipping the checklist altogether because the deadline is close and the draft "looks fine"

A useful compression: Discernment gives you the method for reviewing well; Diligence gives you the reason the review has to happen, and names who answers for it if it didn't.

The two competencies compound rather than operate independently, and the exam's trickiest scenarios exploit the gap between having one without the other. Discernment without Diligence looks like an Associate who knows the five-part checklist cold but, rushed before a deadline, skips running it and ships the draft anyway -- the skill existed, the discipline to apply it under pressure didn't. Diligence without Discernment looks like an Associate who dutifully reviews every output every time, but only checks tone and length, missing a fabricated citation, because they never learned what the checklist actually is -- the discipline existed, the skill to execute it well didn't. Both produce the same visible failure, an error reaching production, but the fix differs: the first needs a workflow gate that forces review regardless of time pressure; the second needs training on what to actually check.

  • A scenario turning on "they didn't realize completeness and relevance are separate failure modes" is testing Discernment -- a skill gap.
  • A scenario turning on "they knew better but skipped the check because the deadline was tight" is testing Diligence -- a discipline gap.
  • Recognizing which one a scenario is testing is itself a form of Discernment, applied one level up to your own review process.

2.1.5 -- Key Concept

Discernment (AI Fluency): judging output against what was requested, what the sources say, and field standards -- the named skill behind the success criteria and checklist in this lesson. Discernment supplies the method for reviewing well; Diligence supplies the reason and timing for why review must happen, and who owns the outcome. Having one doesn't guarantee the other.

2.1.6 Put It Together: Exam Traps for Task Statement 2.1

Task Statement 2.1 questions tend to describe an output that reads impressively — formal tone, confident phrasing, good structure — and then ask what you should do before relying on it. The wrong answers lean on that polish as if it were proof. The right answer always routes back to the same two moves: had success criteria been set, and does the output actually satisfy all five checklist dimensions, especially completeness and relevance, which are the two most often faked by fluent writing.

Hold onto this habit for the rest of the domain: every later skill — spotting hallucinations, fact-checking, deciding on human review, editing, and choosing a format — assumes you've already done this first pass. Evaluation against explicit criteria is the gate everything else in Domain 2 walks through.

ℹ️

Where this shows up on the exam

When a scenario describes a fluent, confident-sounding output, ask yourself: what would the checkable success criteria have been, and does the output actually satisfy all five checklist dimensions — not just the ones that are easiest to notice.

Key Takeaways

  • Fluency is not accuracy — Claude's confident writing style gives no signal about whether the content is correct.
  • Decide what 'good' looks like (explicit, checkable success criteria) before reading or judging the output, not after.
  • The five-part checklist is accuracy, completeness, relevance, consistency, and fitness for audience — each can fail independently.
  • Completeness must be checked against the original request's discrete parts, not against how thorough the answer feels.
  • Relevance means answering the actual question asked, not a fluent-sounding nearby one.
  • The person using Claude stays accountable for the output; self-reported polish is never a substitute for checking it.
  • Discernment (AI Fluency) is the skill of judging output against what was requested, what the sources say, and field standards -- it's the named competency behind the success criteria and checklist taught in this lesson.
  • Discernment answers how to review well; Diligence (named in Lessons 1.3 and 6.4) answers why and when review is required and who owns the outcome -- having one doesn't guarantee the other.

Check Your Understanding

Test what you learned in this lesson.

Q1.A stakeholder reviews a Claude-written project update and says, "This reads really well, let's send it." What should happen before it goes out?

Q2.A request asks Claude to summarize a report, flag any risks, and recommend a next step. The output delivers a strong summary and a clear recommendation but never mentions risks. What checklist dimension has failed?

Q3.Why should success criteria be defined before reading Claude's draft, rather than after?

Q4.A response answers a fluent, well-organized question that is subtly different from the one actually asked. Which checklist dimension catches this?

Q5.An Associate knows the five-part accuracy/completeness checklist thoroughly and has used it well before. Facing a tight deadline today, they skip running it and send the draft as-is, which later turns out to have a dropped requirement. Which competency gap does this best illustrate?

Practice This Lesson

PrepGenAICerts.com is an independent third-party exam-prep platform for the Claude Certified Architect (CCA-F) certification. We are not affiliated with, endorsed by, or acting on behalf of Anthropic PBC.

Note: New premium upgrades are temporarily paused while we resolve an issue with our payment provider. Existing premium members retain full access.