How to Verify an AI Agent Actually Completed a Task
A practical acceptance workflow for Claude Code, Codex, and other agents: define the expected terminal state, read it back from authoritative systems, classify the result, and retry only the missing work.
Contents

You asked Claude Code, Codex, or another agent to organize files, update a spreadsheet, or create business records. It replied “done” and no obvious error appeared. That proves the run ended cleanly; it does not prove the business outcome exists.
A dependable acceptance process starts with the expected terminal state, reads the real result back from the destination systems, and checks both missing and unexpected side effects. Only then should you mark the task complete. When the result is unknown, query first instead of rerunning the entire workflow.
Why a successful tool call is not task completion
Agents can produce several reassuring intermediate signals: a tool returns success, a process exits with code 0, a report is generated, or the final message claims that every step finished. None of these signals answers the questions that matter:
- Was the correct file written to the correct location with the right content?
- Did the spreadsheet or database persist every intended record?
- Were duplicate rows, duplicate orders, extra notifications, or unintended files created?
- Is an asynchronous job still queued, or was a write later rolled back?
- Did the agent only read a record while omitting the required update?
Microsoft’s public ThinkingBox-Bench v1.0 is a synthetic research benchmark, not a production incident rate. It contains 507 executable tasks across five business domains and accepts a run only when checks over the final state, side effects, and designated dialogue properties pass. Its tagged documentation says there is no partial credit. The “partial” and “unverified” labels below are operational acceptance states for your workflow, not benchmark scores.
Define an acceptance contract before the run
The fastest way to verify a task is to state its terminal conditions before the agent starts. Record at least four things:
| Item | What to define | Example |
|---|---|---|
| Required state | Objects, fields, counts, and relationships that must exist | 120 files renamed; 120 unique IDs added to the sheet |
| Forbidden state | Changes and side effects that must not occur | No source files deleted; no second email; no other tab edited |
| Source of truth | The system whose stored state decides acceptance | Destination folder, spreadsheet cells, CRM records, ticket status |
| Retry identity | The key that identifies this business operation | A task- or operation-scoped operation_id or order number; customer IDs, filenames, object IDs, and manifest hashes identify query targets, not this operation by themselves |
Describe the business result, not the agent’s action. “Called the spreadsheet update tool” is not a terminal state. “The destination tab contains 120 rows, every customer ID is unique, and totals reconcile to the source manifest” is.
For irreversible or duplicate-sensitive actions, define the order too. Complete reversible file and record updates first; verify them; only then send email, submit an order, or trigger payment. A failure in the first stage should not force a blind rerun of the irreversible stage.
Preserve enough execution evidence
Keep the smallest evidence set that can connect one run to its writes:
- the original task, allowed scope, and expected terminal state;
- start and finish times, working directory, and target manifest;
- the agent session ID and any returned job ID;
- an
operation_idor another unique business key; - pre-run counts, key fields, versions, or hashes;
- explicit errors, timeouts, permission denials, and skipped steps.
You do not need the agent’s private reasoning to accept the result. You need reproducible inputs, identifiers, scope, and destination state. Redact sensitive fields, and never place API keys, access tokens, or customer secrets in the acceptance log.
Wait for the system’s terminal state, not the agent’s last message
Uploads, bulk imports, report jobs, and third-party writes may continue after the conversation ends. Save the job ID and poll the authoritative system at a reasonable interval until it reports a defined success, failure, cancellation, or timeout state.
Record the last update time and progress while polling. If the destination is eventually consistent, allow a predefined settling window and read it again. Do not classify a just-submitted record as failed merely because it is not immediately visible, and do not wait forever. When the window expires without decisive evidence, the correct status is unverified, not “not executed.”
Read the result back through six evidence layers
1. Confirm that the target object exists and belongs to this run
Check the path, filename, record ID, modification time, and version. An older file with the same name is not proof. A newly created record attached to the wrong customer is not a valid result either.
2. Check content and business invariants
Verify counts, uniqueness, totals, required fields, relationships, and format. For a spreadsheet, compare IDs, row counts, and totals. For files, compare the manifest, sizes, hashes, or sampled contents. For business records, verify status, amount, ownership, and time range.
3. Use the authoritative system of record
The agent’s summary, terminal output, and local cache are supporting evidence only. Acceptance must come from the system that stores the outcome: the file service, actual spreadsheet cells, CRM, ticket system, database, or payment ledger.
4. Verify every required side effect
Some tasks require more than one output: a file, an index update, an associated business record, and a notification. Check each one and make sure they share the same business identity. A missing required side effect means partial completion.
5. Search for effects that should not exist
Acceptance is not only a search for what is present. Look for duplicate rows, duplicate records, extra email, deleted files, out-of-scope edits, and writes to the wrong object. An unexpected side effect deserves its own alert because another retry may amplify it.
6. Reconcile across systems
When a task spans files, sheets, and business applications, first bind every lookup to the current operation_id or task scope, then verify customer and object IDs, the required status or version, and the manifest. Matching an object identifier does not prove that this operation landed. Compare explicit ID sets and list both missing and extra items.
Classify the result with four operational states
| State | Use it when | Next action |
|---|---|---|
| Complete | Every required terminal condition is verified and no unacceptable extra effect exists | Save evidence and close the task |
| Partial | Some outcomes are verified, but specific items or steps are missing | Repair only the missing work |
| Unverified | The system is unavailable, still settling, or evidence cannot determine whether a write occurred | Query or wait; do not blindly retry |
| Unexpected side effect | A duplicate, deletion, out-of-scope change, or wrong-object write exists | Stop automation, contain the damage, and require review |
Unverified does not mean failed. It means you do not yet know. That distinction determines whether the next action is another query or another write.
Decide whether to retry in this order
- Query the current operation first. Scope the system-of-record lookup by
operation_idor order number, then combine filenames, customer or object IDs, and the required status or version. Finding the object alone does not prove that this write occurred. - Identify the completed boundary. List successful objects and missing objects instead of treating the workflow as one opaque success or failure.
- Repair idempotently. If the system guarantees that the same key cannot create a second business result, submit only the missing items. If idempotency is unknown, do not automatically retry an irreversible operation.
- Separate records from notifications and transactions. If the record exists but email failed, resend only the email. If email was sent and record state is unclear, query the record before doing anything else.
- Stop on unexpected effects. Roll back or repair the wrong objects before a human decides whether automation can resume.
A compact rule is: retry only when absence is confirmed; repair only the missing part after a partial write; query before acting when the result is unknown; stop when the system contains the wrong result.
Worked example: files, a sheet, and CRM records
This is a hypothetical example, not a customer incident or a reproduced production log.
An agent must rename 120 files, write 120 rows to a tracking sheet, create 12 summary records in a CRM, and then send one completion email. The read-back shows that all files and rows are correct, only 9 CRM records exist, and the email has already been sent.
The correct status is partial, not complete and not a total failure. A safe recovery is to:
- query the CRM with the
operation_idand all 12 expected business keys; - identify the 9 existing records and create only the missing 3 with deduplication keys;
- reconcile the final set of 12 CRM IDs with the file and sheet summaries;
- avoid renaming the files again, rewriting 120 rows, or resending the email.
If the CRM cannot be queried, mark the result unverified. A full rerun at that point could create duplicate records and another notification.
Copy this acceptance record
| Field | What to record |
|---|---|
| Task and scope | Intended action, allowed destinations, and forbidden changes |
| Expected terminal state | Objects, fields, counts, relationships, and final statuses |
| Execution identity | Session ID, job ID, operation_id, and business keys |
| Authoritative evidence | Destination objects, query time, IDs, and links |
| Missing effects | Required objects or actions that are absent |
| Extra effects | Duplicates, deletions, out-of-scope writes, or extra notifications |
| Acceptance state | Complete, partial, unverified, or unexpected side effect |
| Next action | Close, wait, repair, retry, roll back, or hand off to a person |
Make the record itself auditable. Attach ID lists, a diff, and query timestamps rather than writing only “checked.”
BetterToken can provide the model connection, not outcome acceptance
When connecting Claude Code through BetterToken, follow the current Claude Code setup and connection-check guide. A normal model response without connection or model errors confirms the connection configuration; it does not prove that a file, spreadsheet, or third-party business record reached the intended terminal state.
Use model requests and session identifiers to locate the relevant execution, but base acceptance on the destination files and systems. Model availability, a successful response, or token usage is not evidence of business completion.
Close the task only after three final questions
Before accepting “done,” ask:
- Can I see the expected terminal state in the authoritative system rather than only in the agent’s description?
- Did I check missing results, duplicate writes, and other unexpected side effects?
- If the result is unknown, can I query by a unique business key before deciding to repair or retry?
Close the task only when all three answers are clear. Before the next high-impact agent run, copy the acceptance record above and fill in the terminal state, forbidden state, and retry identity. That small preparation is usually cheaper than untangling a duplicate submission later.