Insights
How to write agent capability claims that hold up
A claim-to-evidence worksheet for describing what an agent does, under which conditions and with which limits.

An agent capability claim should help a prospective user decide whether to try a product for a particular task. When a sentence leaves the task, conditions or evidence unclear, the reader has to supply those details. That is how a narrow success can become a much larger expectation. A practical claim review starts by asking what a reasonable reader would think the product can do after reading the sentence.
This guide uses a fictional project-record assistant to show the process. The assistant reads a brief and proposes updates to an internal project tracker. The examples are writing exercises, not measured results from an operating product. The aim is to produce a claim a team can support and a customer can interpret without a separate conversation.
Begin with the action a user can observe
Suppose a draft homepage says the assistant manages projects autonomously. That sentence could mean scheduling work, assigning people, changing deadlines or simply preparing a summary. A product writer should ask the engineer to demonstrate the exact action the current system performs. In this example, the system reads a project brief and drafts changes to a small set of tracker fields for human review.
The first rewrite could say that it prepares proposed tracker updates from project briefs. This gives the reader a recognizable input and output. A second sentence can explain that a reviewer approves the changes before they are saved. The resulting description is useful even before adding a performance claim because it helps the reader understand the workflow and their own role in it.
Create a claim record
Keep a short record for each meaningful public claim. Include the exact sentence, the product version, the owner of the evidence and the source of that evidence. Add the conditions under which the sentence is true. For the project assistant, those conditions could include a supported tracker, a particular set of fields and briefs written in the format used by the team.
The record should also contain a practical counterexample. Ask what situation would make the claim misleading. A brief might contain a date mentioned as a possibility rather than a commitment. If the assistant treats every date as a confirmed deadline, the team has identified a limitation worth testing and explaining. The counterexample keeps the review attached to user behavior instead of turning into a discussion about adjectives.
Match the evidence to the wording
OpenAI's evaluation guidance calls for tests tailored to the task and its real usage. Applied to claim writing, that means a team should examine the work described in the sentence. Success on an unrelated benchmark would not answer whether this assistant handles the customer's project briefs.
For the fictional assistant, gather examples that resemble the intended inputs and define what a correct proposal contains. Inspect ordinary cases and cases with missing or conflicting information. If the product is described as supporting a particular tracker integration, the evidence should include that integration. A text-only demonstration cannot establish that a field was saved successfully in a remote system.
A team may discover that its strongest evidence supports a smaller claim than it first wrote. That is useful information. It could publish a precise description of the tested function while continuing to investigate broader uses. The wording should move when the evidence moves, and the claim record provides a place to document why a revision was made.
Separate a draft from a completed action
A proposed change, an attempted write and a verified update are different events. A product description should use the verb that matches the event. If a reviewer receives a suggestion, say the system drafts or proposes. If the tool attempts a remote update, describe the attempt accurately. Claiming completion requires a way to establish that the requested state exists after the operation.
Anthropic's agent evaluation discussion treats the final environment state as distinct from the recorded run. This is a useful distinction for a writer reviewing a completion claim. The assistant's final message is one piece of evidence; the actual tracker record is another.
Imagine a network interruption after an update request. The assistant might produce a reassuring message even though the team cannot tell whether the tracker accepted the change. The appropriate product behavior could be an unresolved status and a reconciliation check. Marketing copy that promises every task is completed would hide precisely the situation the operator most needs to understand.
Put the limits beside the promise
A limitation is easier to use when it appears near the relevant claim. If the assistant only supports a defined group of tracker fields, name that scope in the feature description or link directly to the field list. If a person must approve changes, make that part of the normal workflow explanation. A distant disclaimer cannot reliably repair a misleading main sentence.
Try reading the page as a busy customer. Would a person understand that they still need a reviewer? Would they know whether the assistant can create new records or only edit existing ones? Would they expect unsupported attachments to be processed? These questions can reveal missing information that a technically accurate sentence still fails to convey.
Review comparative and numerical language carefully
Words such as faster, better and more accurate imply a comparison. Before using them, identify the baseline, the task and the measurement method. A result from one internal experiment may be worth sharing, but the surrounding text should explain its setting. Avoid turning a small observation into a universal expectation for every customer.
A numerical statement needs a defined denominator. If a report counts accepted proposals, clarify whether rejected inputs, retries and escalations are included. Record the version and the date of the measurement. A product can change after an integration or prompt update, and a claim that once described the system accurately can become stale even when the feature name remains the same.
Let customers test their interpretation
A short user review can expose ambiguities before publication. Give a few intended users the draft sentence and ask what they would expect to happen on a concrete task. Do not explain the intended meaning first. Their answers will show whether the wording communicates the workflow or invites assumptions the product cannot support.
Then show the actual demonstration and ask where it differs from their expectation. A mismatch is a writing problem worth fixing even if the team can defend the original sentence literally. The purpose of product copy is to help people understand an offer, not to create a sentence that survives a narrow reading after someone is disappointed.
Keep a small publication checklist
Before publishing, check that the claim names an observable job, corresponds to the current product and has accessible supporting evidence. Confirm that the relevant limits appear where the reader needs them. Assign someone to revisit the wording after material changes to tools, permissions or supported inputs. A claim record without an owner will quickly become another forgotten document.
Start with one prominent sentence on the homepage or sales deck. Write down its implied task, inspect the strongest evidence and identify one counterexample. Then revise the sentence until a reader can tell what the product does and what they must still decide. That small review is a practical way to bring the public description closer to the system a customer will actually use.
