Claude Code Mods: track context and replay edits without mistaking signals for proof
A practical guide for heavy Claude Code users: choose Token Weather or Replay Theater, load one for a single session, understand its limits, and verify the result with Git, targeted checks, and provider request records.
Contents

Long Claude Code sessions create two different questions: how much context did the latest turn consume, and which file-edit operations did Claude invoke? Anthropic’s playground has a ready-made sample for each question. Token Weather shows main-session context usage; Replay Theater lets you step through file-edit calls from the previous editing turn.
The essential boundary is simple: both are observability tools, not acceptance tests. A context percentage is not your remaining quota, cost, or task completion. An edit in a replay is not proof that you approved it, the tool succeeded, or the final file still contains it.
Pick the Mod that answers your immediate question
| Question | Load first | What it tells you | What it cannot prove |
|---|---|---|---|
| How full is the main context window, and did it grow sharply over recent turns? | Token Weather | Reported context tokens, window size, percentage, and a 12-turn trend | Subscription balance, charge, future requests remaining, or task completion |
Which Edit, Write, or MultiEdit calls occurred in the last editing turn? | Replay Theater | File, tool, local before/after text, and a short diff for each recorded call | Approval, tool success, final disk state, or passing tests |
For diagnosis, load one at a time. Both samples can draw in AbovePrompt. Their READMEs state that this band is shared, so another Mod using it can leave only one visible.
Check the version and trust boundary before loading code
The current sample READMEs require Claude Code 2.1.287 or later and target the terminal interface. Check your client first:
claude --version
These Mods live in an Anthropic DevRel playground. The repository describes its entries as as-is examples without support or a guarantee that they will keep working as Claude Code, the API, and models change. Review the target folder before running it, especially README.md, .claude-plugin/plugin.json, hooks/hooks.json, and the hooks module.
A Mod runs with your user permissions; “it only draws UI” is not a security boundary. For a first trial, prefer --plugin-dir for one session. Closing that Claude Code process ends the trial, and you avoid turning a diagnostic experiment into a persistent installation.
Clone the official samples and validate the one you chose
Clone the playground and move to the Mods directory:
git clone https://github.com/anthropics/claude-code-playground.git
cd claude-code-playground/claude-code/mods
Downloading the repository does not prove that a Mod is structurally valid. Validate the exact folder before starting a session.
For Token Weather:
claude plugin validate ./token-weather
claude --plugin-dir ./token-weather
For Replay Theater:
claude plugin validate ./replay-theater
claude --plugin-dir ./replay-theater
If validate reports an error, stop and fix or restore the named manifest, hooks configuration, or module. Do not continue on the assumption that Claude Code will safely ignore a malformed package. In the new session, /plugin can help confirm what the session loaded. The operational success signal comes next: Token Weather must update after a completed main-loop turn, while Replay Theater needs a completed turn that actually invokes file edits.
Read Token Weather as context telemetry, not a bill
After each main-loop turn, Token Weather calls $.session.usage() and reads tokens, window, and percent from context. It draws one line above the prompt, keeps a chart of the latest 12 readings, and reports how much the latest turn added.
The fields mean:
tokens: the input context over which the latest response was answered, combining uncached, cache-written, and cache-read input tokens;window: the context window for the session’s model;percent:tokens / window.
A 0% reading before the first response is expected because no response has reported usage yet. The display updates once after a turn, not continuously while the turn is running. Subagent turns do not create separate main-loop readings.
Why its percentage can differ from Claude Code’s compaction notice
Token Weather divides by the full context window. Claude Code’s auto-compaction notice is based on a lower compaction point, so the two percentages can differ. The official sample screenshot once showed 81% in Token Weather while the client notice showed 90%; that is an example of two denominators under the sample conditions, not a pair of values you should force to match.
The history bars are relative to the fullest reading currently shown. A low-usage session can therefore display visible bar swings even when the absolute percentage is small. Use the numeric percentage and token count for the absolute reading. The 12-turn history resets when the session starts or the plugin reloads.
What conclusion Token Weather supports
It supports a statement such as: “The main session’s input context rose sharply over the last few turns.” It does not support: “My account has 19% left,” “this turn cost a certain amount,” or “the coding task is complete.” Caching changes billing treatment, but cached input still occupies context. Check balance, charge, and request status in the provider’s own records.
Use Replay Theater to inspect edit attempts
Replay Theater answers “what file-edit calls happened in the previous editing turn?” After loading it, ask Claude to perform a task that actually modifies files and wait for the turn to finish. When the hint appears, open the replay with:
/replay
You can also focus the band with ctrl+x, Tab, then press r. Inside the pane:
| Key | Action |
|---|---|
n | Next step |
p | Previous step |
c or Escape | Close |
Each step shows the file, tool, added and removed line counts, and a short diff. The sample caps a step at 12 displayed lines. For Edit, it compares old_string and new_string, not the complete file, and it does not show file line numbers. For Write, it reads the previous disk contents immediately before the call; files above 400 lines are shown without full line matching.
Why a replayed edit may be absent from the final file
Replay Theater records a call before passing it onward. A change you deny, or a tool call that fails, can still appear. A later call can also overwrite or reverse an earlier one.
The replay lives in memory for the current session. Restarting Claude Code or reloading the plugin removes it. A later turn with no edits keeps the previous replay. Treat it as a record of what Claude attempted, not a snapshot of the final repository.
Accept the work by reading the real repository
Whatever Replay Theater displays, finish with the repository itself:
git status --short
git diff --stat
git diff -- path/to/file
git diff --check
Use git status --short to identify actual additions, modifications, and deletions. Use git diff --stat to spot an unexpectedly broad change. Read the complete diff for important files instead of relying on the 12-line replay excerpt. Run git diff --check for whitespace errors, then execute the smallest test, type check, or build command directly related to the changed files.
Write acceptance criteria as observable outcomes: “the target function was renamed, all references were updated, the relevant unit test passes, and no unrelated file changed.” “The replay showed five green steps” is not an acceptance criterion.
Verify context usage in a provider request record
Token Weather shows how full the session context is; it does not calculate an API bill. With any API provider, match the completed turn by time and model, then check request status, input and output tokens, applicable cache tokens, and the recorded charge.
When BetterToken is the API route for Claude Code, its current site says model, time, token counts, cache usage, final cost, and status stay together in one request record. Use that record to verify actual API usage. This does not mean BetterToken supplies the Mods, stores a complete prompt and response, or automatically accepts the code change. Connection details belong in the BetterToken Claude Code setup guide.
Troubleshoot in the shortest useful order
claude plugin validate fails
Read the exact file and field named by the validator. Confirm that you are in claude-code-playground/claude-code/mods, that a download tool did not rename hidden files, and that the checkout is intact. Revalidate before launching.
Token Weather is missing or stays at 0%
Use the terminal rather than relying on the VS Code chat panel, confirm version 2.1.287 or later, and confirm that the current session loaded the intended Mod. Send a normal request and wait for the main-loop turn to finish. Zero before the first response is normal.
Replay Theater shows no replay hint
Confirm that the turn invoked Edit, Write, or MultiEdit and that the turn has ended. Reading files, answering a question, or only running Bash commands gives Replay Theater no file-edit step to record.
Replay Theater disagrees with git diff
Trust the real files. The edit may have been denied or failed, a later edit may have replaced it, the pane may show only a local excerpt, or session/reload boundaries may have cleared state. Re-evaluate with the full diff and targeted tests.
Two Mods do not appear together
Disable one and start a new session with the other. Because both use AbovePrompt, a missing second band is not by itself evidence that the package is broken.
The minimum reliable loop
Choose Token Weather when your question is context growth. Choose Replay Theater when your question is which file-edit calls occurred. A successful load is only step one, and a visible meter or replay is only step two. Step three is always to inspect the real files, run the smallest relevant validation, and—when usage matters—check the provider request record. That is how you turn process visibility into a verified result.