Claude Code and DeepSeek: Recover from an Input Length Exceeds Maximum Error
Preserve work outside the model, branch on whether /compact succeeds, recover with checkpoints or a fresh session when it fails, and verify the real model, gateway, and compaction limits.
Contents

You add one short instruction in Claude Code and receive status_code=400, Input length 1048609 exceeds the maximum length 1048566. Do not keep retrying, and do not assume that deleting 43 characters will fix it. The safest sequence is to preserve code and task state outside the model, try a focused compaction, fall back to a checkpoint or fresh session if compaction also fails, and then identify whether the model, Claude Code, or an intermediary gateway rejected the request.
The message proves only that one layer considered the input 43 unspecified units over its limit. It does not say whether the unit is tokens, characters, or bytes, and it does not prove that the request reached DeepSeek. Without the actual model ID, base URL, Claude Code version, proxy chain, and raw response headers, you cannot assign this error to a specific model defect.
Preserve completed work before sending another model request
Stop submitting the same long request first. Code is usually already on disk; the fragile part is the task state that exists only in the conversation: what changed, what was verified, and what should happen next.
Use a second terminal to capture the state without asking the over-limit model to summarize it:
claude --version
git status --short
git diff --stat
git diff > claude-context-recovery.patch
git diff --cached > claude-context-recovery-staged.patch
If the project is not in Git, copy the modified files or create an editor snapshot. Then write a short RECOVERY.md manually:
Goal:
Changed files:
Verified:
Still broken:
Key decisions:
Next single action:
Do not reload: full logs, the entire repository, unrelated history
This handoff does not need to be elegant. Its job is to let a clean session recover the task from a few hundred words instead of replaying hundreds of thousands of tokens.
When collecting diagnostics, keep only non-secret facts such as claude --version, the model ID, the base URL domain, proxy names, and the exact error. Do not print API keys, authorization headers, full business prompts, or environment variables that contain credentials.
Distinguish the limit before choosing a fix
Similar-looking length errors can come from different limits.
| Error shape | What it may mean | First action |
|---|---|---|
Input length X exceeds the maximum length Y | A layer enforces a cap on input or request length; the unit may not be tokens | Reduce history and tool output, then identify the rejecting layer |
input length and max_tokens exceed context limit: A + B > C | Input plus reserved output exceeds the model context budget | Compact earlier; consider a smaller output budget only when this is the confirmed condition |
HTTP 413 or request body too large | An HTTP body or gateway byte limit | Inspect uploads, encoding, and proxy body-size settings |
All target providers failed or another generic wrapper | A gateway may have hidden the upstream cause | Inspect the raw upstream attempt or gateway logs |
An early Claude Code user report showed an input-plus-max_tokens 400. Another report said /compact itself failed near the limit because the compaction request needed output space. GitHub #42 and #8136 are independent user reports, not proof of the cause in your environment.
That is why a difference of 43 does not justify removing exactly 43 characters. Claude Code also sends system instructions, conversation history, tool definitions, file contents, and tool results. Intermediaries may count a different unit. Leave meaningful headroom instead of trying to land one unit below the boundary.
Branch A: /compact still succeeds
If slash commands still run, inspect usage and compact with an explicit focus. The current Claude Code documentation defines /context as the context-usage view and /compact [instructions] as replacing history with a summary.
Run:
/context
Then tell compaction what must survive:
/compact Keep the current goal, changed files, verified results, key decisions, unresolved issue, and next action. Drop full logs, repeated file contents, abandoned branches, and unrelated discussion.
After compaction, do not immediately rescan the repository. Test with a narrow, observable request:
Read only RECOVERY.md and src/auth/session.ts. Tell me which function should be changed next. Do not scan other directories and do not edit files.
A successful recovery has four signals:
/contextshows substantially lower usage;- a small request no longer returns 400;
- the model can restate the goal, changed files, and next action;
git diffand test results still match the handoff.
Compaction can omit details. Put durable project rules in the project-root CLAUDE.md, and put critical task state on disk rather than trusting a very long conversation.
Branch B: /compact also returns Conversation too long or the same 400
Do not loop on /compact against the same oversized history. A user report in Claude Code #26317 describes a normal request reaching the limit and /compact then returning Conversation too long. The issue being closed as not planned does not establish that every version and provider path is fixed.
Rewind to before the oversized read or tool output
The current official checkpointing guide says you can run /rewind, or press Esc twice while the prompt is empty, and choose among actions including:
- Restore conversation to rewind the conversation while keeping current code;
- Summarize from here to compress messages after the selected checkpoint;
- Summarize up to here to compress earlier history while preserving later messages.
Choose a checkpoint before the full log paste, whole-file read, or unusually large tool result. Restoring the conversation while retaining code, then issuing a much smaller request, can be safer than trying to summarize the already-overflowing tail.
A summarize action can still require a model call, so this is not guaranteed to work when the upstream rejects the summarization request itself. If rewind or checkpoint summarization gets the same error, move to an empty-context recovery.
Use /clear or a new session when rewind cannot recover
/clear
The official commands reference says /clear starts a new conversation with empty context and keeps the previous conversation available for resumption. It does not undo files on disk or delete the Git diff. Because you already saved patches and RECOVERY.md externally, you do not need the failing model to generate a final summary.
Start the clean session with a constrained instruction:
Read RECOVERY.md, git diff --stat, and the two files listed there. Verify the current state first. Do not scan the whole repository. Propose only the next step and wait for approval before editing.
Do not immediately run claude --continue, claude --resume, or /resume. The official session guide states that resuming restores the full history and tool results. If that history caused the overflow, the next model request can fail again. Keep the old session for later inspection or export, not as the clean recovery path.
Locate the rejecting layer with four low-cost tests
Once work is safe, isolate the cause. Change one variable at a time and reuse the same small read-only task as the control.
| Test | Result | Most useful interpretation |
|---|---|---|
| Same endpoint and model, new empty session, tiny read-only task | Succeeds | Accumulated history or large tool output in the old session is more likely |
| The same tiny task fails in an empty session | Fails | Model mapping, route, request envelope, client version, or server-side cap is more likely |
| Where permitted, compare the official endpoint with the intermediary gateway | Direct succeeds, gateway fails | Inspect gateway limits, conversion, and error aggregation |
/compact fails but a small request after /clear succeeds | Succeeds | The compaction request exceeded the real window, or the client assumed the wrong window for the custom model |
Capture non-secret routing values:
printf 'BASE_URL=%s\nMODEL=%s\nOPUS=%s\nSONNET=%s\nHAIKU=%s\nSUBAGENT=%s\n' \
"${ANTHROPIC_BASE_URL:-<unset>}" \
"${ANTHROPIC_MODEL:-<unset>}" \
"${ANTHROPIC_DEFAULT_OPUS_MODEL:-<unset>}" \
"${ANTHROPIC_DEFAULT_SONNET_MODEL:-<unset>}" \
"${ANTHROPIC_DEFAULT_HAIKU_MODEL:-<unset>}" \
"${CLAUDE_CODE_SUBAGENT_MODEL:-<unset>}"
The command intentionally does not print ANTHROPIC_AUTH_TOKEN or any other secret. Also record whether the request passes through Claude Code Router, a switcher, a corporate gateway, a reverse proxy, or multi-provider fallback. A request can be rejected before it reaches the model.
A September 2026 report in Claude Code Router #1799 says a gateway replaced an upstream context error with All target providers failed, preventing the client from seeing the specific cause. The reporter’s local patch is not a universal production fix, but the case shows why support logs should retain the upstream status, body, and request ID.
Verify the actual DeepSeek route, not just a [1m] label
As of September 30, 2026, the client-side environment-variable example on DeepSeek’s official Claude Code integration page sets:
ANTHROPIC_MODEL,ANTHROPIC_DEFAULT_OPUS_MODEL, andANTHROPIC_DEFAULT_SONNET_MODELtodeepseek-flash[1m];ANTHROPIC_DEFAULT_HAIKU_MODELandCLAUDE_CODE_SUBAGENT_MODELtodeepseek-flash;CLAUDE_CODE_AUTO_COMPACT_WINDOW=786432.
The same page separately says that, when a request passes Claude-style model names, DeepSeek maps names starting with claude-opus to deepseek-v4-pro, and names starting with claude-haiku or claude-sonnet to deepseek-flash. That service-side mapping rule is a separate fact from the explicit values in the client example; they should not be collapsed into one claim about the actual route.
Neither section establishes which model or context window handled this customer’s request. Check available request or usage records, the raw model field, the selected provider or route, and the upstream request ID instead of inferring from the UI or an older tutorial.
A [1m] label also does not guarantee that every layer accepts one million tokens. The client, compatibility API, gateway, fallback provider, and final model can each impose a separate limit, and reserved output still consumes context. A clean-session control plus the raw upstream response is stronger evidence than a label alone.
Compact earlier instead of inflating the window
DeepSeek’s current official example includes:
export CLAUDE_CODE_AUTO_COMPACT_WINDOW="786432"
This is a client-side threshold for earlier compaction, not a way to enlarge the provider’s hard limit. The current Claude Code environment-variable reference says the value must be a plain integer, is bounded by the model’s context window, and overrides /autocompact, launch flags, and settings.
If the actual model or gateway allows less, set a lower verified value instead of copying 786432. If you need more safety margin, CLAUDE_AUTOCOMPACT_PCT_OVERRIDE can lower the percentage at which auto-compaction triggers; it cannot raise the threshold.
Do not treat CLAUDE_CODE_MAX_CONTEXT_TOKENS as an “unlock context” switch. It tells Claude Code the verified real window of a custom or unrecognized model route. Setting it above the server’s actual cap delays compaction and makes an upstream 400 more likely.
Everyday practices are often more effective than increasing a number:
- filter large logs locally with
grep,rg,tail, or a script before sharing excerpts; - read large files by function or line range instead of loading the whole file;
- use
/clearbetween unrelated tasks; - delegate broad research or repository scans to a subagent so the main conversation receives only a summary;
- run a focused
/compactbefore a new phase, not at the final few tokens; - avoid repeating the same schemas, full build logs, or complete diffs.
Retest with one small, real task
After recovery or a configuration change:
- start a clean session;
- read only
RECOVERY.mdand one small file; - request a short, bounded answer;
- confirm the 400 is gone;
- add files and tool calls gradually.
If even the small request fails, stop trimming the old conversation and investigate endpoint configuration, model mapping, proxy behavior, and request format. If only the long session fails, focus on compaction timing, tool-output size, and session boundaries.
A useful support report includes the UTC timestamp, Claude Code version, exact model ID, base URL domain, proxy chain, raw status and error body, request ID, clean-versus-old-session result, and whether /compact reproduced the error. Omit keys, authorization headers, and full business content.
Frequently asked questions
Why did restarting Claude Code not help?
Restarting does not necessarily create empty history. --continue, --resume, and the resume picker restore the old conversation. Use a genuinely clean session and a tiny control request to test whether history is the cause.
Will lowering output tokens fix it?
Only when the error explicitly says input plus max_tokens exceeds a combined context budget. If input alone exceeds an independent limit, lowering output does not guarantee a fix.
Will switching gateways always solve it?
No. A different route can help when the current gateway enforces a smaller body limit or masks upstream errors. No gateway can exceed the final model’s hard context limit.
Does /clear delete code?
No. It clears conversation context, not files already written to disk. Still verify with git status, a patch, a commit, or editor history before clearing.
Why not delete exactly 43 characters?
The error does not identify the unit, and the request contains hidden system content, tool definitions, and history. Creating substantial headroom is more reliable than aiming for exactly 43 fewer visible characters.
The recovery sequence to remember
For Input length exceeds maximum, use this order: save the diff and a manual handoff in another terminal → run /context → try focused /compact → if compaction fails, /rewind to before the oversized output → if that still fails, /clear or start a new session → use a tiny control request to locate the model, client, or gateway limit → set auto-compaction below the verified server window.
This process does not depend on guessing what the number 43 means, and it does not claim that a proxy can bypass a model’s hard limit. It protects the work first, then turns a vague 400 into a repeatable layer-by-layer diagnosis.