cruxlevel

Claude Associate · study module 2/7

Output Evaluation and Validation

21% of the exam · ~75 min read

Cheat sheet — the night-before version

  • Heaviest domain: 21% — judgment questions, not recall
  • Verification effort scales with the cost of being wrong — never uniform
  • Accuracy audit ≠ completeness audit: check what's there AND what's missing
  • "Are you sure?" / self-confirmation / regenerate-and-compare = NOT verification
  • A citation you can't open and confirm yourself doesn't exist
  • Web-search citations lower risk but still need a click-through: real page ≠ supports the claim
  • Internal contradiction (summary says 12%, table says 8%) = stop and reconcile against source
  • Sycophancy: prompts that telegraph the desired answer tend to get it — re-ask neutrally
  • Human review is mandatory before external, legal/compliance, financial, or irreversible outputs
  • Refine for audience IN the conversation with directional feedback, not from scratch
  • Format follows use: inline = discuss, artifact = iterate/share a deliverable, structured data (table/CSV) = feed another system
  • Fluency, confidence, and length are zero evidence of accuracy

What this domain tests

21% of the exam — the heaviest domain by a clear margin, and the one the exam guide leads with in its own sample questions. The bet behind the weighting: an Associate's value isn't producing text with Claude, it's knowing whether that text is accurate, complete, unbiased, and safe to send — and what to do when it isn't.

Character: this is a judgment domain, not a recall domain. First-hand reviews of the Associate exam describe it as non-technical, centered on "hallucination detection and organizational fit," with answers that are less clear-cut than the technical exams — several options look defensible, and you must pick the most appropriate action. Expect workplace scenarios: a summary bound for the compliance team, a stat headed into a board deck, a colleague who says "it has citations, so we're done." The official sample question for this domain is exactly that shape: Claude confidently cites a subsection number; the right answer is "verify the cited subsection against the official regulation text," and the distractors are self-reported confidence and cosmetic rewording.

Six official task objectives. Every one appears below.

Theory outline

1. Risk-tiered verification — the organizing principle

The domain's spine: verification effort scales with the cost of being wrong. Not everything gets the same scrutiny, and both extremes are wrong answers on the exam.

StakesTypical outputsVerification standard
LowBrainstorms, internal rough drafts, throwaway reformattingSkim for obvious nonsense
MediumTeam reports, blog drafts, internal analyses others rely onCheck key claims, all numbers, all names/dates
HighCustomer-facing content, legal/compliance/regulatory, financial figures, board material, anything published or irreversibleVerify every load-bearing fact against a primary source; human sign-off required

Two symmetric distractor families to recognize:

  • Over-verification: "read every word of every output with two reviewers." Doesn't survive real volume; the exam treats it as a failure of judgment, not diligence.
  • Under-verification: "Claude expressed high confidence, so send it." Confidence is a register, not a signal.

The correct-answer shape is almost always a checklist keyed to the costly failure modes — check the pricing, the legal claims, the numbers; skim the prose.

2. Evaluating for accuracy AND completeness

The objective says "accuracy and completeness" — two different audits, and the exam distinguishes them.

Accuracy audit — is what's there correct?

  • Numbers, dates, names, titles, citations: verify each against the source or a primary reference.
  • Claims about your documents: trace them back to the actual passage. If Claude says "the contract allows termination with 30 days' notice," find that clause.
  • Quotes: a generated quote is fabricated until the person actually said it or approved it.

Completeness audit — is anything load-bearing missing?

  • Summaries silently drop things. A summary of 200 survey responses into four themes can be accurate on all four and still bury a fifth theme that mattered.
  • Techniques: compare against a known checklist of what should appear (agenda items, required sections, all business units); probe the source yourself for topics the output never mentions; ask targeted questions about sections you already know well and compare answers to ground truth.
  • Absence claims are the hardest case. "The document does not contain a non-compete clause" is a confident negative that a missed synonym can falsify. Negatives with consequences get your own manual keyword pass over the source.

3. Identifying hallucinations

A hallucination is fluent, plausible, fabricated content — invented statistics, citations, quotes, names, events, legal provisions — delivered in the same confident register as correct output. That register-match is the whole problem: nothing about the text tells you it's wrong.

Working tells the exam rewards:

TellWhy it's suspect
Precise, unsourced specifics ("$4.7B market by 2027")Exact numbers with no provenance are the classic fabrication signature
Citations you can't locate (404s, no search hits)Treat as nonexistent; the claim comes out with it
Very specific details about niche/recent topicsSparse training coverage is where fabrication concentrates
Perfectly convenient evidence (the ideal quote, the study that proves your exact point)Plausibility optimized for your request, not reality
Confident answers past the knowledge cutoffThe model can't know it; grounding (web search) is required, not better prompting

Non-fixes the exam loves as distractors: asking "are you sure?" (models confidently confirm their own fabrications), regenerating (re-rolls the dice), softening a dead citation to "studies show" (launders the fabrication), rewording to sound more formal (cosmetic), and running it three times to see if answers agree (consistency ≠ correctness — a model can be reliably wrong).

Legitimate risk-reducers per Anthropic's own guidance: allow an explicit out ("say 'I don't know' if the information isn't in the document"), require quoted evidence from the source before conclusions, ask Claude to flag its own uncertainty and assumptions, and ground fast-moving topics with web search so claims come with inspectable, current sources. These reduce hallucination rates; none replaces verification for high-stakes output.

4. Identifying inconsistencies

Distinct from hallucination: the output disagreeing with itself or with its own source.

  • Internal contradiction: the executive summary says revenue grew 12%; the table below says 8%. At least one is wrong, and you don't know which — the only correct move is reconciling both against the source before the document moves anywhere.
  • Output vs. source drift: the summary says "most customers were satisfied" when the data shows 55/45. Directionally shaped, technically defensible, materially misleading.
  • Cross-run variation: two colleagues ask the same question, get different answers, and conclude "it's broken." Generation is probabilistic; variation is expected. The professional response is verifying facts that matter against sources — not repeated asking, and not treating either run as an oracle.

5. Identifying biases

Three flavors show up in scenarios:

  • Sycophancy — the model picks up your framing and agrees with it. If you ask "what's the strongest argument against my plan?" while making your enthusiasm obvious, the counterargument comes back weak. First suspect: your own prompt. Fix: re-ask neutrally, in a fresh chat, without revealing your position — or have a colleague who disagrees run the same question.
  • Framing/anchoring in the prompt — "summarize why this vendor is the right choice" produces advocacy, not analysis. Evaluation questions reward noticing that the prompt predetermined the output's slant.
  • Systematic skew in outputs — e.g., deal summaries that consistently overstate customer enthusiasm, or candidate screens that echo patterns from training data. A consistent bias is a prompt-and-process problem: add a calibration rule (claims require quoted evidence; default to neutral), then spot-check outputs against sources to confirm the fix. For decisions about people (hiring, performance), bias risk is one of the standing triggers for human review — that thread continues in the Governance domain.

6. Fact-checking and validation techniques

The toolbox, ranked by evidentiary value:

  1. Primary sources. The regulation text itself, the contract itself, the dataset itself. The official sample answer for this domain is literally "verify the cited subsection against the official regulation text."
  2. Click-through citation verification. When Claude uses web search, answers arrive with linked sources. That lowers risk and creates an audit trail — but a citation can point to a real page that doesn't say what the claim attributes to it (citation drift). For load-bearing claims: open the source, find the supporting passage. "It has links" is the beginning of verification, not the end.
  3. Source-anchored spot-checks for document summaries: trace key claims to specific sections; probe sections you know well and compare with ground truth.
  4. Independent recomputation for anything numeric: totals, percentages, date math. Arithmetic is cheap to check and embarrassing to publish wrong.
  5. Domain-expert review where you can't judge correctness yourself — you can verify a citation exists without being able to verify the legal reasoning is sound.

Explicitly not validation: the model's own confidence, self-confirmation, output fluency or length, regeneration agreement, and "it matches what I expected" (expectation-matching is how wrong outputs get through).

7. When human review or additional verification is required

The exam wants you to route, not just verify. Standing triggers where a human gate is non-negotiable:

  • External audiences — customers, press, regulators, partners. Nothing AI-drafted leaves the building unreviewed.
  • Legal, compliance, medical, financial content — reviewed by someone qualified in that domain, not just any human.
  • Irreversible or costly actions — payments, contract commitments, public statements, deletions.
  • Decisions about people — hiring, firing, performance, credit. Both a bias risk and, in many jurisdictions, a regulatory line: AI recommends, a human decides.
  • Claude expresses uncertainty, or you can't verify a load-bearing claim — uncertainty flagged by the model is a routing signal, not a formality to strip out of the draft.
  • The output surprised you — an answer that contradicts what you expected deserves a second look in both directions: your expectation might be stale, or the output might be wrong.

The matching principle from workflow design: place the gate before the irreversible step, and scale it with error cost — a step whose errors are cheap or caught downstream can run reviewed-by-exception; a step whose errors are silent and costly gets a mandatory checkpoint however boring it is.

8. Editing, adapting, refining, and comparing outputs for the audience

Accurate-but-unusable is a real failure mode, and this objective covers fixing it.

  • Refine in-conversation with directional feedback. Claude keeps context: "keep the structure; cut to one page; replace jargon with plain terms a non-financial executive knows" beats "make it better" (arbitrary changes) and beats starting over (loses everything established).
  • Audience adaptation is specification: reading level, vocabulary, length, what leads, what's cut. One analysis legitimately becomes three deliverables — a one-paragraph executive brief, a detailed appendix, a customer-safe version — by re-rendering the same verified content per audience. Verify the facts once; adapt the presentation per audience; then re-check that adaptation didn't distort meaning (aggressive simplification can turn "correlates with" into "causes").
  • Comparing outputs: when you have two candidate drafts (two prompts, two models, before/after an edit), compare against explicit criteria — accuracy against source, fit to audience, completeness against requirements — not against vibes. Have Claude critique a draft against a rubric if useful, but the pick is yours: model self-assessment is an input, never the verdict.
  • The reset move: after many rounds of contradictory steering, outputs degrade. Consolidate everything learned into one clean prompt and start fresh — the exception that proves the iterate-in-place rule.

9. Organizing information and selecting output formats

The objective names three surfaces: artifacts, inline, structured data. Format follows what happens to the output next.

FormatWhat it isChoose when
Inline (chat response)Prose/bullets in the conversationYou're discussing, exploring, or the answer is consumed once, in place
ArtifactSubstantial content in a dedicated pane — documents, reports, code, mini-apps — iterable and shareable as its own objectThe output is a deliverable you'll revise across turns, reference repeatedly, or hand to someone else; keeps the work product separate from the discussion about it
Structured data (table, CSV, JSON)Machine-consumable rows/fields with consistent structureThe output feeds another system or workflow step — spreadsheets, trackers, databases — or must be filtered, sorted, or compared row-by-row

Curation is the other half of the objective: don't ship the raw transcript of an exploration. Organize the finding — pull the decision and its rationale into the deliverable, keep the supporting detail retrievable, cut the dead ends. A correct answer that nobody can locate inside twelve paragraphs fails the intended audience just as surely as a wrong one. Structured formats also make validation easier: a table of claims with a source column is auditable in a way flowing prose never is.

Worked problems

Problem 1 — the contradicting report

An analyst has Claude draft a quarterly business review from an uploaded financials workbook. The executive summary states revenue grew 12% quarter-over-quarter; a table later in the same document shows 8%. The deck is due to the leadership team in an hour. What should the analyst do?

A. Delete the table — the prose summary is the part executives read B. Ask Claude which number is correct and update the document to match its answer C. Recompute the growth figure from the uploaded workbook, correct whichever instance is wrong, then scan the document for other figures that disagree with the source D. Regenerate the whole report and use the new version if its numbers are internally consistent

Approach: An internal contradiction means at least one claim is wrong, and the document itself can't tell you which. The only arbiter is the source data. A prepared candidate also generalizes: one caught inconsistency is evidence the generation process produced errors, so the other load-bearing numbers now need a pass too.

Why the others fail: A hides the contradiction instead of resolving it — the prose number might be the wrong one. B asks the system that produced the error to adjudicate it; Claude may confidently pick either. D re-rolls the dice, and internal consistency in the new version is not evidence of correctness — a regenerated report can be consistently wrong.

Answer: C

Problem 2 — the four-theme summary

A product manager has Claude distill 250 open-ended survey responses into key themes. Claude returns four themes, each with representative quotes. Before presenting them as "what our customers said," what is the MOST important additional check?

A. Confirm the quotes are well-chosen and persuasive B. Check whether a meaningful share of responses falls outside the four themes — for example by asking Claude what didn't fit, then sampling the raw responses yourself C. Ask Claude to rate its confidence in each theme and drop any below 80% D. Verify the themes are ordered from most to least frequent

Approach: The quotes make the accuracy of what's present easy to spot-check — the open risk is completeness. Summarization silently drops content, and "what our customers said" is a completeness claim. The check must probe what the output doesn't contain, which requires going back to the raw data, not interrogating the summary.

Why the others fail: A checks persuasiveness, not coverage — polish is not evidence. C is self-reported confidence, the distractor the official sample question exists to punish. D is cosmetic ordering; a perfectly ordered list can still omit the fifth theme that mattered most.

Answer: B

Problem 3 — "it has citations"

For a client-facing regulatory brief, an associate uses Claude with web search enabled. The answer includes linked sources for its claims. A teammate says: "It's citing real websites — that means it's verified. Send it." What is the accurate assessment?

A. Correct — web search grounds answers in real sources, which is what verification means B. Incorrect — web search results are less reliable than training knowledge, so the links should be removed C. Partially right: linked sources lower risk and give you an audit trail, but for a client-facing brief you must open the load-bearing sources and confirm each actually supports the specific claim attributed to it D. Incorrect — citations should be verified only when they look suspicious

Approach: Grounding via web search is a genuine improvement: claims arrive with inspectable, current sources. But a real link can be attached to a claim the page doesn't actually make — the claim-to-source mapping is itself generated. Stakes are high (client-facing, regulatory), so load-bearing citations get a click-through and a passage check.

Why the others fail: A confuses having sources with checking them — the links are the audit trail, not the audit. B is backwards; grounding in current sources is precisely the fix for staleness. D fails because fabricated or drifted citations look exactly like sound ones — "looks suspicious" is not a detection method, which is the whole reason systematic verification exists.

Answer: C

Problem 4 — routing to human review

A junior operations associate uses Claude to produce four outputs in one afternoon. Which one REQUIRES qualified human review before use, rather than the associate's own spot-check?

A. An internal brainstorm of possible names for a team dashboard B. A reformatted version of last week's meeting notes for the team wiki C. A draft response to a customer's formal complaint that references the company's warranty obligations D. A first-pass categorization of internal IT tickets that the IT team will triage anyway

Approach: Run the standing triggers: external audience? legal/compliance content? irreversible? caught downstream? Option C hits two triggers at once — it leaves the building AND makes claims about legal obligations — and warranty misstatements to a complaining customer are exactly the silent, costly error class. "Qualified" matters: this wants someone who can judge the warranty language, not just any second pair of eyes.

Why the others fail: A is low-stakes and internal — skim territory; escalating it is the over-verification trap. B is internal with cheap, visible errors. D has a built-in downstream catch: the IT team triages every ticket anyway, so miscategorizations get corrected in the normal flow.

Answer: C

Problem 5 — one analysis, three audiences

A finance analyst has verified a Claude-drafted cost analysis and now needs versions for (1) the CFO, who wants one paragraph, (2) the finance team, who needs the full detail, and (3) an all-hands slide for non-finance staff. What is the BEST way to produce the three versions?

A. Send the full analysis to all three audiences so nothing is lost B. In the same conversation, ask Claude to re-render the verified analysis for each audience with explicit specs (length, vocabulary, what leads), then re-check each version for meaning drift C. Start three fresh chats and rewrite the analysis from scratch for each audience to avoid cross-contamination D. Ask Claude which audience matters most and produce only that version

Approach: Facts are verified once; presentation adapts per audience. The conversation already holds the verified analysis — in-context re-rendering with concrete audience specs is the efficient move. The prepared candidate adds the final beat: adaptation is a new generation step, so each version gets a quick check that simplification didn't distort meaning (hedged findings becoming flat assertions is the classic drift).

Why the others fail: A ignores the audience objective entirely — the CFO asked for a paragraph; sending forty is a usability failure. C discards the verified content and context, re-introducing error risk three times over for no benefit — "cross-contamination" isn't a concern between renderings of the same approved analysis. D invents a prioritization decision nobody asked the model to make; all three audiences have a stated need.

Answer: B

Problem 6 — picking the output surface

A procurement lead asks Claude to compare seven software vendors on price, security certifications, support tiers, and contract terms. The team will import the comparison into their existing tracking spreadsheet, filter it, and re-sort it during negotiations. Which output format should the lead request?

A. A well-organized inline prose comparison, one paragraph per vendor B. An artifact containing a polished narrative report with an executive summary C. A structured table (exportable as CSV) with one row per vendor and one column per criterion D. A bulleted list of each vendor's overall strengths and weaknesses

Approach: Format follows the output's next step. The stated destination is another system (a tracking spreadsheet) plus row-level operations (filter, sort, compare). That's the structured-data profile, verbatim. Bonus: a criteria-per-column table is also the easiest format to validate — every cell is a discrete, checkable claim.

Why the others fail: A and D lock the data inside prose — someone has to manually re-extract every value into the spreadsheet, adding a transcription-error step. B is the right surface for a deliverable someone reads and iterates on, not data feeding a system; a polished narrative artifact still can't be filtered or sorted.

Answer: C

Exam traps

  • Self-confirmation as verification. "Ask Claude to double-check / rate its confidence / say if it's sure" — never the answer for anything that matters. The official sample question's wrong options are built from exactly this.
  • Fluency as evidence. Polished, confident, well-structured output gets zero credibility points. Distractors that reason from tone are always wrong.
  • Consistency as correctness. "Ran it three times, same answer" or "the regenerated version agrees" — a model can be reliably wrong. Agreement between generations is not agreement with reality.
  • Citation presence as citation verification. Links from web search are an audit trail to walk, not a verdict. Real page ≠ supports the claim.
  • Laundering instead of removing. Softening an unverifiable citation to "studies show," or a fabricated quote to "the CEO has indicated" — the exam treats disguising a fabrication as worse than the fabrication.
  • Uniform scrutiny in either direction. "Review everything exhaustively" and "it's internal, skip review" both lose to risk-tiered checking. Watch for the option that names the specific high-risk items to verify.
  • Accuracy-only audits. Verifying every claim that's present while never asking what's missing. Completeness questions hide behind accuracy framing — "the summary checked out" is not "the summary is complete."
  • Fixing usability problems with regeneration or model switches when the real gap is audience specification, and fixing systematic bias by rewriting outputs by hand instead of fixing the prompt and spot-checking the fix.
  • Format mismatch. Prose where data needs to flow into a system; a chat reply where the team needs a shareable, iterable deliverable. When the stem tells you what happens to the output next, it's telling you the format answer.

Prefer being taught? This guide exists as an adaptive tutor course inside Claude Code — your agent teaches it, quizzes your weak spots, and tracks progress. How it works · included in membership (founding: $0)

Now prove it: the Output Evaluation and Validation quiz

Original practice questions on exactly this domain — instant explanations, domain score, and your wrong-answer notebook at the end.

Take the quiz →

Next module: Product and Model Selection