Insights
Build a small evaluation casebook for an agent
Practical cases for checking saved results, missing evidence, conflicting sources and failed actions.

A small evaluation casebook can give an agent team a shared answer to a practical question: what should happen on this task? The casebook is a set of concrete examples with enough context to run, inspect and discuss them. It is useful before launch because it turns vague concerns into cases that people can examine together.
This guide develops a fictional document assistant. A user asks it to update a project handoff document from approved notes, save the result and report anything it could not resolve. The cases below are suggested starting points, not a validated test suite. A real team would need examples from its intended work and agreement about the behaviors it considers acceptable.
Write the task before the expected answer
Start with one complete request a user might make. For example: update the handoff document with the confirmed delivery date from the latest approved project note, preserving the rest of the document. The case needs the starting document, the available notes and a definition of which note is approved. Without those details, two reviewers can imagine different tasks while discussing the same sentence.
Next, describe the expected state. The saved document should contain the confirmed date in the relevant field and preserve unrelated content. If no approved date exists, the assistant should leave the field unchanged and identify the missing information. This definition gives the team a basis for judging the result without requiring one exact wording for the assistant's final message.
Use a compact case record
Give each case a stable identifier and a short title. Store its inputs, permitted tools, expected outcome and escalation rule. Add a note explaining why the case belongs in the set. The identifier lets people refer to the example in a bug report or release discussion, while the rationale helps future maintainers understand whether the case still matters.
For the document assistant, a record might identify the source folder, the initial document version and the allowed output location. It should also state whether the agent may overwrite a file or must create a revision. These operational details affect the outcome. A beautifully written update saved to the wrong folder can still fail the user's task.
Case one: a routine successful update
Create an ordinary case with one approved note and one clear date. Keep the surrounding document realistic enough to reveal accidental edits. After the run, inspect the saved file and compare the relevant field. Check that names, headings and unrelated sections remain intact. Also inspect whether the assistant's final message accurately describes the change.
The routine case establishes that the harness itself works. If a test runner cannot find the output or compare the correct file version, a failure may come from the test setup. Resolve that before drawing conclusions about the agent. Save the input and observed output so a teammate can reproduce the investigation.
Case two: the evidence is missing
Remove the approved note while leaving a draft that mentions a possible date. The desired behavior is to preserve the existing value and explain that the confirmed date is unavailable. The test asks whether the assistant recognizes the difference between a plausible answer and authorized evidence. It should also check whether the unresolved item is easy for the user to find.
An escalation rule can make the expected behavior concrete. The assistant may ask the user for the approved note or flag the field for review. It should not silently choose the draft date. Reviewers should agree on that rule before seeing the run; otherwise, a persuasive answer can change the grading standard after the fact.
Case three: current sources disagree
Supply two approved notes with different dates and no clear precedence rule. The assistant should surface the conflict rather than invent a resolution. The case record can require it to identify both sources and leave the document unchanged. This tests a different skill from handling missing evidence: recognizing that available information does not support a single answer.
A follow-up variant can add an explicit precedence rule, such as an approved revision number. Now the task is resolvable, and the assistant should apply the stated rule. Keeping these cases separate makes it possible to distinguish unnecessary escalation from unjustified confidence. Both behaviors matter to a user trying to finish work.
Case four: saving fails
Simulate a tool failure when the assistant tries to save the updated document. The final state should remain unchanged, and the assistant should report that the save did not complete. Check the actual artifact instead of accepting the final message as proof. Anthropic's evaluation article distinguishes the recorded run from the outcome in the environment, a distinction that matters directly in this case.
A related variant can make the save result uncertain. Perhaps the request times out after it was sent. The correct next action depends on the tool's guarantees. A reconciliation read may establish whether the change occurred. If it cannot, the assistant should preserve uncertainty and avoid an unsafe repeated write. The casebook should record that expected behavior explicitly.
Choose checks that match the question
Some checks can be exact. A file exists at a specified path, a date field matches the approved source or unrelated text remains unchanged. Other checks need judgment, such as whether an escalation message gives the user enough information to act. Separate these checks in the record so a failure does not collapse several different problems into a single score.
For a judgment check, write a short rubric with examples of acceptable and unacceptable behavior. Have two people inspect a few outputs independently, then discuss disagreements. If they cannot agree, improve the rubric before automating the judgment. A model-based grader can be useful, but it should not conceal a standard the team itself has not made clear.
Keep the cases representative
OpenAI recommends evaluations tied to real task distributions. In this casebook, that means balancing ordinary work with meaningful failure situations. A set containing only unusual traps could misrepresent how the assistant performs on the work people actually give it. A set containing only clean examples would miss the moments where users need help most.
Collect candidate cases from permitted samples of intended work, with sensitive information removed or replaced appropriately. Ask users where they currently pause, double-check or seek clarification. Those moments can reveal valuable cases. Keep a note about where each example came from so the team can understand the coverage and its gaps.
Use the casebook during change review
Run the established cases when a material change affects the workflow, then inspect differences by case. A new prompt may improve the routine update while making conflict handling worse. An aggregate score can hide that tradeoff. The casebook makes the change visible and gives the team a concrete decision to discuss.
Keep a separate group of new examples for checking whether improvements extend beyond familiar cases. Record the product version, tool behavior and run date with each result. Start small enough that the team will actually read the failures. A useful first milestone is a casebook that two colleagues can run and interpret consistently, with unresolved questions clearly documented for the next review.
