Output Evaluation and Validation
21% of examCritically evaluate Claude's output for accuracy, completeness, hallucinations, inconsistencies, and bias, apply fact-checking and validation techniques scaled to the stakes, decide when human review is required, and edit and format the result for its audience and destination.
6
task statements
13
concepts
78
practice questions
Domain Mastery
Evaluate output accuracy and completeness
Judging Claude's output against explicit success criteria for accuracy, completeness, relevance, consistency, and audience fit before it is used.
Knowledge of
- Defining success criteria up front -- what 'good' looks like -- before judging an output, per Anthropic's success-criteria guidance
- The practical checklist: accuracy, completeness, relevance, consistency, and fitness for audience
- Why polish and confident phrasing are not evidence of correctness
- The difference between an output that merely sounds complete and one that actually addresses every requested part
- Discernment as the AI Fluency competency for judging output against what was asked for, what the sources actually say, and the standards of your field -- and how it differs from Diligence
Skills in
- Writing explicit, checkable success criteria (e.g., word limits, required citations, prohibited recommendations) before evaluating an output
- Checking an output against each practical-checklist dimension rather than forming a single gut impression
- Confirming that every discrete part of a multi-part request was addressed, not just the most prominent one
- Resisting the pull of confident, well-written prose as a substitute for verification
- Recognizing whether an exam scenario is testing the skill of reviewing (Discernment) or the discipline of deciding to review at all (Diligence)
Concepts
Defining Success Criteria Before Evaluating Output
✎CoreDecide what 'good' looks like before evaluating an output, not after
The Accuracy, Completeness, Relevance, Consistency, and Audience-Fit Checklist
✎CoreThe five-part checklist is accuracy, completeness, relevance, consistency, and fitness for audience
Discernment: The AI Fluency Competency Behind This Entire Task Statement
✎CoreDiscernment is the AI Fluency competency for judging output against what was requested, what the sources say, and field standards -- it's the named skill behind the success criteria and five-part checklist taught in this task statement
Recognize hallucinations, inconsistencies, and bias in Claude's output
Identifying confident-but-false content, internal contradictions, and skewed framing in Claude's output, and applying grounding techniques that reduce -- but do not eliminate -- the risk.
Knowledge of
- Hallucination as confident, plausible-looking, but false or fabricated content -- invented statistics, citations, sources, or quotes
- The three places hallucinations concentrate: specific-looking details, the edge of the model's knowledge, and long outputs where one fabrication hides among correct content
- Inconsistency (internal contradiction) and bias (skewed framing, unrepresentative examples, assumptions about people) as distinct categories of 'unexpected' output
- Capability hallucination -- Claude claiming to have taken an external action (sent an email, saved a file) that it has no tool to actually perform
- Grounding, the 'I don't know' exit, requesting traceable citations/quotes, and cross-checking as the techniques that reduce -- not eliminate -- hallucination risk
- Why self-reported model confidence is not a reliable accuracy signal, and why grounding/RAG reduces but does not eliminate hallucination risk
Skills in
- Spotting specific-looking numbers, dates, citations, and quotes as the highest-risk spots for fabrication
- Distinguishing a hallucination from an internal inconsistency from a bias issue so the correct fix follows
- Recognizing a claimed external action (an email 'sent,' a file 'saved') as unverified until independently confirmed, since claude.ai has no tool to perform it
- Prompting Claude to ground answers in provided source material and to say 'I don't know' rather than guess
- Requesting traceable citations/quotes and cross-checking specific claims against an authoritative source
- Not treating a model's stated confidence, or the presence of grounding/RAG, as proof of accuracy
Concepts
Hallucination: Confident but False or Fabricated Content
✎CoreHallucination = confident, plausible-looking output that is false or fabricated (invented stats, citations, sources, quotes)
Inconsistencies, Bias, and Hallucination-Reduction Techniques
✓AdvancedHallucination (fabrication), inconsistency (internal contradiction), and bias (skewed framing) are three distinct failure types requiring different fixes
Capability Hallucination: Claiming an Action It Never Took
✎CoreCapability hallucination is Claude claiming to have taken an external action (emailed a file, saved a document) it never actually took
Apply fact-checking and validation techniques scaled to stakes
Matching the validation technique -- source verification, recomputation, human review -- to how the output will be used and how high the stakes are.
Knowledge of
- Validation as the deliberate step of confirming an output before it is used, distinct from making it look more polished
- The stakes-scaled validation approaches: verifying names/numbers/dates/citations, opening source text for cited regulation/policy, recomputing numeric totals, confirming a summary adds nothing beyond its source, and human review for customer/legal/compliance-facing output
- Quote-grounding (asking Claude to extract supporting quotes before drawing conclusions) and best-of-N comparison (re-running a request and comparing outputs) as additional grounding and validation techniques
- The exam's canonical case: a cited regulation subsection must be verified against the official text before it is shared
Skills in
- Selecting a validation technique that matches the type of claim (citation, number, summary) and how the output will be used
- Verifying a cited subsection, statute, or policy against the actual source text rather than trusting Claude's citation
- Recomputing or spot-checking numeric analysis against source data
- Asking Claude to quote supporting passages before analyzing a long document, so reasoning and errors stay visible
- Re-running a request and comparing outputs (best-of-N) to build confidence where runs agree and flag soft spots where they diverge
- Refusing to substitute reformatting, or 'it's just internal,' reasoning for genuine verification
Concepts
Fact-Checking and Validation Scaled to Stakes
✎CoreValidation is a deliberate confirmation step, distinct from improving an output's polish or formatting
Quote-Grounding and Best-of-N Comparison
✓AdvancedQuote-grounding: for long documents, have Claude extract supporting quotes before analyzing, so conclusions trace back to a specific line rather than a trust-me summary
Determine when human review is required
Recognizing the specific conditions -- high stakes, external publication, unverifiable claims, sensitive data, or scope beyond the Associate role -- that require escalating an output to human review.
Knowledge of
- The conditions that require human review: high-stakes (legal/financial/medical/compliance/safety/reputational) content, external publication without further checks, unverifiable factual claims, regulated or sensitive data or policy decisions, and tasks beyond Associate scope
- Human-in-the-loop as a feature of responsible use, not a failure of the tool or a sign Claude is inadequate
- The judgment skill of matching review level to actual stakes, rather than escalating everything or nothing
Skills in
- Screening an output against the human-review trigger conditions before it is used or shared
- Escalating technical or out-of-scope tasks to an Architect/Developer rather than attempting them beyond Associate scope
- Calibrating review intensity to stakes instead of defaulting to blanket escalation or blanket automation
Concepts
Edit, adapt, refine, and compare Claude's output
Treating Claude's output as a draft to be corrected, reframed for its audience, polished, and compared against alternatives before it is finalized.
Knowledge of
- Raw Claude output as a draft, not a finished deliverable
- The four standard moves: editing for correctness, adapting for a different audience, refining tone/structure, and comparing alternative versions
- Requesting multiple variations or an audience-specific rewrite as standard techniques for improving a draft
- Code execution as the mechanism that produces a genuinely verified, computed numeric result -- distinct from a plausible-looking prose figure
Skills in
- Editing an output for correctness before treating it as final
- Reframing the same content for a different audience (executives vs. engineers, a Slack message vs. a formal client email)
- Asking Claude for two or three variations and selecting the strongest rather than accepting the first draft
- Requesting a rewrite for a different reader instead of manually rewriting from scratch
- Asking Claude to compute a number via code execution rather than write it in prose, when the number must be right
- Recognizing that a deterministic, re-runnable calculation still depends on correct code, so the result should be checkable, not blindly trusted
Concepts
Editing, Adapting, Refining, and Comparing Output
✎CoreTreat Claude's first response as a draft, not a finished deliverable
Code Execution: Computing Numbers Instead of Writing Them
✎CoreWhen a number must be right, have Claude compute it via code execution rather than generate it in prose
Organize information and choose the right output format
Matching the output surface -- inline response, Artifact, or structured data -- to how a deliverable will be used, and validating structured output's content, not just its shape.
Knowledge of
- Inline responses as the fit for short answers, quick edits, and conversational back-and-forth
- Artifacts as a separate, editable window suited to substantial, self-contained deliverables that will be refined, reused, or shared
- Structured data (tables, JSON) as the fit when output feeds a spreadsheet, form, or downstream system and needs a predictable shape
- Structured/well-formed output still needing content validation -- valid shape is not valid content
Skills in
- Choosing inline, Artifact, or structured output based on how the deliverable will actually be used, not by default
- Recognizing when a substantial, iterate-and-share deliverable belongs in an Artifact rather than buried in chat
- Validating the content of structured/JSON output rather than trusting it because it parses correctly
- Avoiding a one-format-fits-all habit across different output needs
Concepts
Inline, Artifacts, and Structured Output as Distinct Surfaces
✎CoreInline suits short, conversational answers; Artifacts suit substantial, iterate-and-share deliverables; structured data suits output feeding another system
Matching Format to Use, Not Forcing One Format
✓AdvancedFormat selection is a judgment call tied to how the deliverable will be used, not a fixed habit