PrepGenAICerts

Screening a Whole Use Case: The Four Delegation Criteria

Core

Distinguish appropriate from inappropriate use cases · Difficulty 3/5

0%
delegationuse-case-screeningai-fluencyaccountabilitygovernance

Explanation

Everything so far in this task statement has been about a single request bumping against a data or policy restriction. There's a related but distinct skill: classifying an entire use case -- not "can I upload this file" but "should this whole workflow be handed to Claude at all." The tool for that job is the same one AI Fluency's Delegation competency uses to size up a single workflow step, just aimed at a bigger target: how reversible it is, what an error would cost, whether it needs a distinctly human touch, and who's accountable.

The Four Screening Questions

  • Reversibility -- If something goes wrong, is there a window to catch it and fix it before real damage is done?
  • Consequence of error -- Setting reversibility aside, what's the actual cost if the output turns out wrong -- money, time, someone's trust, someone's job?
  • Human creativity or empathy -- Is there a layer of judgment, relationship, or genuine care in this task that a model can't actually provide, regardless of how polished its output reads?
  • Accountability -- Is there a specific person who has to answer for how this turns out, and are they actually in a position to exercise that responsibility over something AI generated?

None of these four is a solo gatekeeper. Failing one doesn't automatically sink the use case, and passing all four doesn't automatically clear it either -- what matters is how they interact for *this* specific case.

Find the One Doing the Real Work

Run all four, then ask a follow-up question: which ONE of them, if it changed, would actually move the classification? That single criterion is the one carrying the weight of the decision, and pointing to it explicitly is what makes a classification defensible instead of a gut call.

Two short contrasts make the interaction concrete:

  • A low-stakes, fully reversible task can still call for a human, when the human-element criterion is the one doing the real work. A retirement send-off card can be resent or corrected with zero real cost, and on paper that reads as low-risk -- but it's inherently about a relationship, and a machine-drafted sentiment nobody bothered to personalize reads as exactly that the moment it's noticed. Here, the human element is what's actually load-bearing, not the stakes.
  • A high-stakes task can still clear the bar, when a properly designed gate hands accountability back to a person. A board-facing quarterly summary carries real consequences if it's wrong -- but having a named reviewer sign off before distribution turns what feels like an irreversible risk into something gated and checkable. Here, accountability is what's load-bearing, and the gate is what satisfies it, not any inherent safety in the task itself.

Three Classifications

  • Fully appropriate -- the outcome is undoable, the downside is small, and nothing about it calls for a distinctly human touch. Hand it off and review it the normal way.
  • Appropriate with human review -- genuinely worth AI assistance, but the stakes or who's on the hook for the outcome mean a person has to be built into the process. That safeguard has to be spelled out concretely (see the companion concept on the who/what/when gate), not just asserted.
  • Inappropriate -- the downside, the fact that it can't be undone, or the need for a distinctly human touch is severe enough that AI shouldn't be doing this task at all. State the reason, and name the person who owns it instead.

Worked Example: Layoff Notifications

A manager asks Claude to draft individual layoff notification letters for a round of reductions, to be sent directly to affected employees. Screen it: reversibility is low (once sent, the message is out, and a callous or incorrect letter can't be un-sent from someone's memory); consequence of error is severe (a wrong severance detail or an impersonal tone causes real harm and legal exposure); the human-element requirement is high (this is exactly the kind of moment that requires empathy and direct human ownership); accountability is diffuse if AI drafts and auto-sends. The load-bearing criterion is the human-element requirement -- even with perfect accuracy, a person needs to deliver this news. Classification: inappropriate for AI to draft-and-send unsupervised. The appropriate human role is HR/the manager, who may use Claude only to prepare talking points or check severance figures for accuracy -- not to generate the notification the employee receives.

Worked Example: Auto-Approving Low-Value Expense Reimbursements

A finance team wants Claude to auto-approve expense reimbursements under $50 that match a receipt and a pre-approved category, with no human touch. Screen it: reversibility is high (a wrongly approved $50 item is trivially clawed back or absorbed); consequence of error is low (bounded dollar amount, no downstream decision rides on it); there's no human-creativity or empathy requirement at all -- this is pattern-matching against a rule; accountability is preserved because the rule itself was set and owned by a person, and the AI is just executing it consistently. The load-bearing criterion is consequence of error, and it's low enough that the other three don't need to compensate for anything. Classification: fully appropriate to delegate, with normal periodic audit review rather than per-transaction human sign-off.

Worked Example: AI-Drafted Performance Review Ratings

A manager wants Claude to draft numeric performance ratings for a team based on project notes, to go into the HR system. Screen it: reversibility is moderate (a rating can be corrected later, but it may have already shaped a compensation or promotion decision by then); consequence of error is high (ratings materially affect pay and career trajectory); the human-element requirement is significant (fair evaluation of a person's work is exactly the kind of judgment call that needs a human's context, not just what's captured in project notes); accountability is the sharpest issue -- if a rating is challenged, someone has to be able to explain and stand behind it. Multiple criteria are severe here, but accountability is the one that, if resolved, actually changes the outcome: a rating a manager reviews, edits, and personally signs is defensible; the same rating auto-submitted is not. Classification: appropriate with human review -- Claude may draft a first-pass rating from notes, but the manager reviews, adjusts, and owns the final number before it enters the HR system.

Common exam traps

  • Treating any single failed criterion as automatically disqualifying, or any single passed criterion as automatically clearing -- the exam rewards naming which criterion is load-bearing for the specific scenario, not mechanically tallying pass/fail across all four.
  • Classifying a task as 'inappropriate' when a defined human gate would actually make it appropriate-with-review -- don't default to blocking a use case when a gate would preserve its value (the same 'find the compliant path' instinct from earlier in this task statement).
  • Calling a task 'fully appropriate' just because it's low-stakes in one dimension (e.g., reversible) while ignoring that the human-element criterion is the one actually driving the classification.

Key Takeaways

  • The same four criteria that screen a single delegated workflow step also screen an entire use case: reversibility, consequence of error, need for human creativity/empathy, and accountability
  • The four criteria interact rather than acting as independent pass/fail gates -- run all four, then identify the single one actually carrying the decision for this scenario
  • A low-stakes, reversible task can still need a person involved if the human-element criterion is doing the real work (e.g., a message that carries a relationship)
  • A task can be high-consequence yet appropriate, if a defined gate restores accountability and reversibility (e.g., a named reviewer sign-off)
  • Three classifications: fully appropriate (hand off, review normally), appropriate with human review (spell out the safeguard concretely), inappropriate (state the reason, and name who should own it instead)
  • Naming the load-bearing criterion is what makes a use-case classification defensible to a reviewer, rather than a gut call

Glossary Terms

Related Concepts

PrepGenAICerts.com is an independent third-party exam-prep platform for the Claude Certified Architect (CCA-F) certification. We are not affiliated with, endorsed by, or acting on behalf of Anthropic PBC.

Note: New premium upgrades are temporarily paused while we resolve an issue with our payment provider. Existing premium members retain full access.