The Gate Is the Classification: Defining Who/What/When
CoreDistinguish appropriate from inappropriate use cases · Difficulty 3/5
Explanation
Learners misapply the middle classification more than any other. The error isn't choosing 'appropriate with human review' when it doesn't fit -- it's stopping there, as if naming the category were the same as building the safeguard. Calling something 'a human-reviewed process' describes a bucket, not a control. Nothing is actually protected until the review itself is spelled out.
Three Elements a Real Gate Needs
Every workable gate answers three concrete questions:
- WHO reviews -- name the accountable role, not an available body. 'The engagement partner' or 'the benefits administrator' counts; 'someone from the team' does not, because no one in particular owns the failure if it slips through.
- WHAT gets checked -- the precise failure the review is guarding against. Not 'give it a once-over,' but the actual thing that could go wrong: a miscalculated figure, a discriminatory pattern, language that misstates a legal obligation.
- WHEN it happens -- before the output goes anywhere, never after. A check performed once the harm is already done is an autopsy, not a safeguard.
Testing Two Statements Against Each Other
Take a lending scenario and compare two ways of describing the same intended control:
| Statement | Does it function as a gate? |
|---|---|
| "There will be human oversight of the approval recommendations." | No -- no named reviewer, no stated failure mode, no point in the process specified |
| "The underwriting supervisor checks every AI-flagged denial for disparate-impact indicators before the applicant is notified." | Yes -- reviewer (underwriting supervisor), failure mode (disparate-impact indicators), timing (before notification) are each spelled out |
The second version can be audited later: ask the supervisor what they looked for, confirm they looked before anyone was told no. The first version gives an auditor nothing to check, because nothing specific was ever promised.
Applying This to an Earlier Example
Go back to the performance-rating scenario discussed under use-case screening, classified as needing human review. The label alone leaves the job unfinished -- what finishes it is naming the reviewer as the employee's own manager (not HR broadly, not a peer), stating that they're checking whether the draft rating reflects the employee's full body of work rather than just what made it into scattered notes, and fixing the timing as before the number is entered into the HR system rather than as a later audit. Framed that way, someone reviewing the process afterward has something concrete to verify. Framed only as 'a manager looks at it,' they don't.
Why the Exam Cares
A scenario that asks you to classify a use case is really asking two things back to back: pick the right category, and then, if that category involves human review, describe the safeguard in a form that could actually be checked later. A choice that just repeats the category ('maintain suitable human oversight') without naming a reviewer, a failure mode, and a moment in the workflow is a wrong answer wearing the shape of a right one.
Common exam traps
- Accepting a restated label ('ensure human oversight,' 'have someone review it') as a complete answer -- these describe the classification, not the safeguard.
- Naming a reviewer without naming what they're checking for -- a review with no defined failure mode can't be verified as having worked.
- Placing the check after the output already shipped -- a review that happens once the outcome has taken effect is a postmortem, not a safeguard, and doesn't satisfy the human-review requirement.
Key Takeaways
- 'Appropriate with human review' is unfinished until the safeguard names a reviewer, a failure mode, and a point in the workflow
- Reviewer = a specific accountable role, not whoever is available; failure mode = the precise risk being checked for; timing = before the output is used, never after
- A restated label like 'ensure human oversight' is not itself a safeguard -- it has no named reviewer, no stated risk, and no timing, so nothing about it can be audited
- A checkable example: 'the underwriting supervisor checks every AI-flagged denial for disparate-impact indicators before the applicant is notified'
- If a use case's safeguard can't be described with a named reviewer, a specific check, and a timing, the use case is not actually ready to run, no matter how confident the classification sounds
Glossary Terms
Related Concepts
Screening a Whole Use Case: The Four Delegation Criteria
The same four criteria that screen a single delegated workflow step also screen an entire use case: reversibility, consequence of error, need for human creativity/empathy, and accountability
Anthropic's Usage Policy as the Outer Boundary
The AUP is the outer boundary of acceptable use; organizational policy layers additional rules on top
Bias, Transparency, and Accountability in AI-Assisted Work
AI output can reflect or amplify bias; review deliberately, especially in decisions about people