How to Verify an AI Agent Actually Completed a Task

A practical acceptance workflow for Claude Code, Codex, and other agents: define the expected terminal state, read it back from authoritative systems, classify the result, and retry only the missing work.

Contents
How to Verify an AI Agent Actually Completed a Task

You asked Claude Code, Codex, or another agent to organize files, update a spreadsheet, or create business records. It replied “done” and no obvious error appeared. That proves the run ended cleanly; it does not prove the business outcome exists.

A dependable acceptance process starts with the expected terminal state, reads the real result back from the destination systems, and checks both missing and unexpected side effects. Only then should you mark the task complete. When the result is unknown, query first instead of rerunning the entire workflow.

Why a successful tool call is not task completion

Agents can produce several reassuring intermediate signals: a tool returns success, a process exits with code 0, a report is generated, or the final message claims that every step finished. None of these signals answers the questions that matter:

  • Was the correct file written to the correct location with the right content?
  • Did the spreadsheet or database persist every intended record?
  • Were duplicate rows, duplicate orders, extra notifications, or unintended files created?
  • Is an asynchronous job still queued, or was a write later rolled back?
  • Did the agent only read a record while omitting the required update?

Microsoft’s public ThinkingBox-Bench v1.0 is a synthetic research benchmark, not a production incident rate. It contains 507 executable tasks across five business domains and accepts a run only when checks over the final state, side effects, and designated dialogue properties pass. Its tagged documentation says there is no partial credit. The “partial” and “unverified” labels below are operational acceptance states for your workflow, not benchmark scores.

Define an acceptance contract before the run

The fastest way to verify a task is to state its terminal conditions before the agent starts. Record at least four things:

ItemWhat to defineExample
Required stateObjects, fields, counts, and relationships that must exist120 files renamed; 120 unique IDs added to the sheet
Forbidden stateChanges and side effects that must not occurNo source files deleted; no second email; no other tab edited
Source of truthThe system whose stored state decides acceptanceDestination folder, spreadsheet cells, CRM records, ticket status
Retry identityThe key that identifies this business operationA task- or operation-scoped operation_id or order number; customer IDs, filenames, object IDs, and manifest hashes identify query targets, not this operation by themselves

Describe the business result, not the agent’s action. “Called the spreadsheet update tool” is not a terminal state. “The destination tab contains 120 rows, every customer ID is unique, and totals reconcile to the source manifest” is.

For irreversible or duplicate-sensitive actions, define the order too. Complete reversible file and record updates first; verify them; only then send email, submit an order, or trigger payment. A failure in the first stage should not force a blind rerun of the irreversible stage.

Preserve enough execution evidence

Keep the smallest evidence set that can connect one run to its writes:

  • the original task, allowed scope, and expected terminal state;
  • start and finish times, working directory, and target manifest;
  • the agent session ID and any returned job ID;
  • an operation_id or another unique business key;
  • pre-run counts, key fields, versions, or hashes;
  • explicit errors, timeouts, permission denials, and skipped steps.

You do not need the agent’s private reasoning to accept the result. You need reproducible inputs, identifiers, scope, and destination state. Redact sensitive fields, and never place API keys, access tokens, or customer secrets in the acceptance log.

Wait for the system’s terminal state, not the agent’s last message

Uploads, bulk imports, report jobs, and third-party writes may continue after the conversation ends. Save the job ID and poll the authoritative system at a reasonable interval until it reports a defined success, failure, cancellation, or timeout state.

Record the last update time and progress while polling. If the destination is eventually consistent, allow a predefined settling window and read it again. Do not classify a just-submitted record as failed merely because it is not immediately visible, and do not wait forever. When the window expires without decisive evidence, the correct status is unverified, not “not executed.”

Read the result back through six evidence layers

1. Confirm that the target object exists and belongs to this run

Check the path, filename, record ID, modification time, and version. An older file with the same name is not proof. A newly created record attached to the wrong customer is not a valid result either.

2. Check content and business invariants

Verify counts, uniqueness, totals, required fields, relationships, and format. For a spreadsheet, compare IDs, row counts, and totals. For files, compare the manifest, sizes, hashes, or sampled contents. For business records, verify status, amount, ownership, and time range.

3. Use the authoritative system of record

The agent’s summary, terminal output, and local cache are supporting evidence only. Acceptance must come from the system that stores the outcome: the file service, actual spreadsheet cells, CRM, ticket system, database, or payment ledger.

4. Verify every required side effect

Some tasks require more than one output: a file, an index update, an associated business record, and a notification. Check each one and make sure they share the same business identity. A missing required side effect means partial completion.

5. Search for effects that should not exist

Acceptance is not only a search for what is present. Look for duplicate rows, duplicate records, extra email, deleted files, out-of-scope edits, and writes to the wrong object. An unexpected side effect deserves its own alert because another retry may amplify it.

6. Reconcile across systems

When a task spans files, sheets, and business applications, first bind every lookup to the current operation_id or task scope, then verify customer and object IDs, the required status or version, and the manifest. Matching an object identifier does not prove that this operation landed. Compare explicit ID sets and list both missing and extra items.

Classify the result with four operational states

StateUse it whenNext action
CompleteEvery required terminal condition is verified and no unacceptable extra effect existsSave evidence and close the task
PartialSome outcomes are verified, but specific items or steps are missingRepair only the missing work
UnverifiedThe system is unavailable, still settling, or evidence cannot determine whether a write occurredQuery or wait; do not blindly retry
Unexpected side effectA duplicate, deletion, out-of-scope change, or wrong-object write existsStop automation, contain the damage, and require review

Unverified does not mean failed. It means you do not yet know. That distinction determines whether the next action is another query or another write.

Decide whether to retry in this order

  1. Query the current operation first. Scope the system-of-record lookup by operation_id or order number, then combine filenames, customer or object IDs, and the required status or version. Finding the object alone does not prove that this write occurred.
  2. Identify the completed boundary. List successful objects and missing objects instead of treating the workflow as one opaque success or failure.
  3. Repair idempotently. If the system guarantees that the same key cannot create a second business result, submit only the missing items. If idempotency is unknown, do not automatically retry an irreversible operation.
  4. Separate records from notifications and transactions. If the record exists but email failed, resend only the email. If email was sent and record state is unclear, query the record before doing anything else.
  5. Stop on unexpected effects. Roll back or repair the wrong objects before a human decides whether automation can resume.

A compact rule is: retry only when absence is confirmed; repair only the missing part after a partial write; query before acting when the result is unknown; stop when the system contains the wrong result.

Worked example: files, a sheet, and CRM records

This is a hypothetical example, not a customer incident or a reproduced production log.

An agent must rename 120 files, write 120 rows to a tracking sheet, create 12 summary records in a CRM, and then send one completion email. The read-back shows that all files and rows are correct, only 9 CRM records exist, and the email has already been sent.

The correct status is partial, not complete and not a total failure. A safe recovery is to:

  • query the CRM with the operation_id and all 12 expected business keys;
  • identify the 9 existing records and create only the missing 3 with deduplication keys;
  • reconcile the final set of 12 CRM IDs with the file and sheet summaries;
  • avoid renaming the files again, rewriting 120 rows, or resending the email.

If the CRM cannot be queried, mark the result unverified. A full rerun at that point could create duplicate records and another notification.

Copy this acceptance record

FieldWhat to record
Task and scopeIntended action, allowed destinations, and forbidden changes
Expected terminal stateObjects, fields, counts, relationships, and final statuses
Execution identitySession ID, job ID, operation_id, and business keys
Authoritative evidenceDestination objects, query time, IDs, and links
Missing effectsRequired objects or actions that are absent
Extra effectsDuplicates, deletions, out-of-scope writes, or extra notifications
Acceptance stateComplete, partial, unverified, or unexpected side effect
Next actionClose, wait, repair, retry, roll back, or hand off to a person

Make the record itself auditable. Attach ID lists, a diff, and query timestamps rather than writing only “checked.”

BetterToken can provide the model connection, not outcome acceptance

When connecting Claude Code through BetterToken, follow the current Claude Code setup and connection-check guide. A normal model response without connection or model errors confirms the connection configuration; it does not prove that a file, spreadsheet, or third-party business record reached the intended terminal state.

Use model requests and session identifiers to locate the relevant execution, but base acceptance on the destination files and systems. Model availability, a successful response, or token usage is not evidence of business completion.

Close the task only after three final questions

Before accepting “done,” ask:

  1. Can I see the expected terminal state in the authoritative system rather than only in the agent’s description?
  2. Did I check missing results, duplicate writes, and other unexpected side effects?
  3. If the result is unknown, can I query by a unique business key before deciding to repair or retry?

Close the task only when all three answers are clear. Before the next high-impact agent run, copy the acceptance record above and fill in the terminal state, forbidden state, and retry identity. That small preparation is usually cheaper than untangling a duplicate submission later.

Ready to optimize your LLM workflow?

Join thousands of developers building faster, smarter, and more cost-effective AI applications with BetterToken.

Get Started for Free