DeepSeek V4 Flash and V4.1 Flash for Coding: Official Benchmarks, max_tokens, Pricing, and Tools
A practical guide that separates DeepSeek V4 Flash, the 0731 release, and the current V4.1 Flash. It consolidates official GPQA, SWE-bench, Terminal-Bench, DeepSWE, NL2Repo, and HumanEval results; explains the difference between a 1M context window and max_tokens; and shows pricing, Python usage, coding-tool setup, and a reproducible way to measure real cost per accepted task.
Contents

People searching for “DeepSeek V4 Flash benchmark official” usually need more than a claim that the model is fast or inexpensive. They want to know which scores DeepSeek actually published, which test conditions produced them, and whether those results say anything useful about day-to-day coding. A search for “deepseek v4 max_tokens” is even more specific: how large is the context window, how much can one response generate, and what value should be sent in an API request?
The first distinction matters most: DeepSeek V4 Flash, V4 Flash 0731, and the current DeepSeek V4.1 Flash are different releases. The original V4 Flash used a 284B-parameter MoE architecture with roughly 13B parameters active per token. V4.1 Flash moved to a 552B-parameter MoE backbone, with about 8B active during prefill and 16B during decoding. Combining the old architecture, the current model ID, and scores from several releases creates a neat-looking specification sheet that cannot be reproduced.
Release status as of September 18, 2026: DeepSeek’s current API model ID is
deepseek-flash. The old V4 Flash and Vision aliases are in a compatibility-routing period and may temporarily resolve to V4.1. For production evaluations, save the model ID, date, reasoning level, and provider instead of relying on a legacy alias alone.
Key specifications
| Specification | DeepSeek V4 Flash / 0731 | DeepSeek V4.1 Flash |
|---|---|---|
| Status | Historical release; legacy aliases may be rerouted | Current Flash release |
| Context window | 1,000,000 tokens | 1,000,000 tokens |
| Architecture | 284B MoE, about 13B active | 552B MoE; about 8B active in prefill and 16B in decode |
| Long-output guidance | Local high/max configurations recommended a 384K maximum output length | Current official API ceiling: up to 384K; local benchmark reproduction recommends max_tokens >= 256K |
| Input modality | Text | Text and images |
| Reasoning control | Earlier high/max thinking configurations | API supports controls such as reasoning_effort |
| Current API model ID | deepseek-v4-flash is a legacy-style name | deepseek-flash |
Neither “384K” nor “256K” should be read as a universal hard limit for every hosted API. They are recommendations for different releases and execution setups. A provider, SDK, gateway, or account policy may expose a lower cap. Check the current endpoint’s model metadata before designing a production workflow around a number copied from an older article.
What does max_tokens actually control?
max_tokens caps the number of tokens the response may generate. It does not resize the model’s context window. The context budget generally includes the input, conversation history, tool results, and space reserved for the answer. A model can support a 1M-token context without making a 256K- or 384K-token response sensible for every request.
Reasonable starting points are:
max_tokens=4096for code explanations, one-function fixes, short SQL, and configuration debugging.max_tokens=16384for multi-file change plans, longer test reports, and migration guidance.max_tokens=32768or more for repository analysis, long agent traces, or large code generation—only after checking the endpoint cap and verifying that the task benefits from that much output.
A value that is too small can cut off a patch, test result, or conclusion. A value that is unnecessarily large expands the worst-case cost and latency. For coding agents, a bounded budget plus repeated read-edit-test cycles is usually safer than asking for an entire repository solution in one response.
Also distinguish visible output tokens, reasoning tokens, and billable tokens. Providers may report them differently. Use the usage fields and the invoice from the endpoint you actually call.
Official benchmarks: match the release, mode, and harness
Official scores answer a narrow question: what did the model achieve under a documented evaluation setup? They do not guarantee the same result on your repository. The tables below keep releases separate and retain the original benchmark names so that scores from different harnesses or reasoning settings are not collapsed into one synthetic rating.
Representative results from the original V4 Flash model card
| Benchmark | Score | How to read it |
|---|---|---|
| GPQA Diamond (Pass@1) | 88.1 | Reported for a high/max reasoning setup |
| LiveCodeBench (Pass@1) | 91.6 | Code-generation result |
| SWE-bench Verified (Resolved) | 79.0 | Real repository issue resolution |
| Terminal-Bench 2.0 (Acc) | 56.9 | Terminal-agent tasks |
| HumanEval Base (Pass@1) | 69.5 | Base-model result; not directly comparable with Max mode |
These numbers come from DeepSeek’s model card, not one uniform independent rerun. HumanEval Base in particular uses a different setting from the high-reasoning rows, so subtracting 69.5 from 91.6 would not measure a meaningful capability gap. The table is best used to see which capability areas were evaluated.
Official family comparison: 0731, V4 Pro, and V4.1 Flash
| Benchmark | V4 Flash 0731 | V4 Pro | V4.1 Flash |
|---|---|---|---|
| GPQA Diamond | 89.9 | 92.4 | 90.9 |
| Terminal-Bench 2.1 | 82.7 | 87.9 | 90.6 |
| Terminal-Bench 4.0 | 7.0 | 12.4 | 31.2 |
| DeepSWE v1.1 | 54.4 | 62.7 | 74.2 |
| NL2Repo-Bench | 54.2 | 61.5 | 64.0 |
The useful conclusion is not that every metric improved without exception. It is that V4.1 Flash made substantial gains in terminal agents, repository-level software engineering, and natural-language-to-repository work. Terminal-Bench 4.0 is also much harder than 2.1, so those two rows must not be treated as interchangeable versions of one score.
V4.1 Flash versus frontier models in the official table
| Benchmark | V4.1 Flash | GPT-5.6 Sol | Opus-5.0 | GLM-5.3 |
|---|---|---|---|---|
| GPQA Diamond | 90.9 | 94.1 | 93.4 | 88.1 |
| Terminal-Bench 2.1 | 90.6 | 88.8 | 89.1 | 88.2 |
| Terminal-Bench 4.0 | 31.2 | 39.9 | 51.8 | 37.9 |
| DeepSWE v1.1 | 74.2 | 73.0 | 74.0 | 66.9 |
| NL2Repo-Bench | 64.0 | 56.8 | 75.3 | 58.0 |
This comparison was published in the DeepSeek V4.1 model card, so it should be treated as vendor-reported data, not a fully neutral league table. Compare models within the same row and harness, then validate the pattern with independent sources and your own tasks. V4.1 is strong on Terminal-Bench 2.1 and DeepSWE v1.1, but it is not the top model in every row, especially Terminal-Bench 4.0 and NL2Repo-Bench.
HumanEval remains useful as a quick function-level code-generation check, but it is too narrow for modern coding agents. SWE-bench, DeepSWE, Terminal-Bench, and NL2Repo are closer to actual work because the model must inspect a repository, use tools, edit files, run tests, and recover from failures.
How to read benchmark results correctly
- Verify the harness version. Terminal-Bench 2.0, 2.1, and 4.0 represent different task sets and difficulty levels.
- Record the reasoning level and output budget.
low,high, andmaxcan produce very different success rates, latency, and token use. - Convert request price into cost per successful task. A cheap model that needs three retries may cost more than a higher-priced model that finishes once.
- Prioritize a rerun on your own repository. Dependency installation, test duration, tool permissions, file count, and code conventions all change agent performance.
Official benchmarks are useful for shortlisting models. They are not a substitute for routing, procurement, or production acceptance tests.
Pricing: low token rates do not guarantee low task cost
As of September 18, 2026, DeepSeek’s official baseline rates for deepseek-flash were:
| Meter | Peak price | Off-peak price |
|---|---|---|
| Cache-hit input | $0.006 / 1M tokens | $0.003 / 1M tokens |
| Cache-miss input | $0.30 / 1M tokens | $0.15 / 1M tokens |
| Output | $1.20 / 1M tokens | $0.60 / 1M tokens |
These are DeepSeek’s official baseline rates, not a guaranteed BetterToken checkout price. BetterToken’s catalog is updated from its live pricing API, and access groups, cache rules, minimum billable units, or retry policies may differ. Check the live rate, save the quote date, and calculate from actual usage:
request_cost =
cache_hit_input / 1_000_000 * cache_hit_rate
+ cache_miss_input / 1_000_000 * cache_miss_rate
+ output_tokens / 1_000_000 * output_rate
cost_per_accepted_task = sum(request_costs) / accepted_tasks
For coding, the most useful metrics are usually first-pass test success, average retries, total tokens per accepted patch, time to green tests, and human rework minutes. “Price per million tokens” is only one input.
Artificial Analysis offers another task-level perspective by combining the scale of its evaluation suite, output volume, speed, and estimated cost. Its methodology is not identical to a provider invoice, but the approach is more informative than comparing list prices alone.
Python API example
The example below uses the OpenAI-compatible SDK, BetterToken’s Base URL, and the current model ID. A 4K output budget is enough to confirm connectivity and inspect basic code quality. Do not upload an entire repository merely to “use” the 1M context window.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["BETTERTOKEN_API_KEY"],
base_url="https://www.bettertoken.ai/v1",
)
response = client.chat.completions.create(
model="deepseek-flash",
messages=[
{"role": "user", "content": "Review this Python function, find why the test fails, and propose the smallest fix. Do not refactor unrelated code."}
],
max_tokens=4096,
reasoning_effort="low",
)
print(response.choices[0].message.content)
reasoning_effort="low" is a practical default for fast iteration. Compare high or max when a task involves a difficult regression, cross-file dependencies, or several rounds of tool use. If the installed SDK does not expose this field directly, pass it through the SDK’s extra request arguments or follow the gateway’s current documentation.
Cursor, Cline, Aider, OpenCode, and other coding tools
- Cursor / Cline / Aider / OpenCode: choose an OpenAI-compatible provider, set the Base URL to
https://www.bettertoken.ai/v1, usedeepseek-flashas the model, and keep the API key in an environment variable or the tool’s secure secret store. - Codex and external agents: this works only when the tool supports a custom OpenAI-compatible endpoint. Check whether it appends
/v1automatically to avoid a duplicated path. - Claude Code: it natively speaks the Anthropic protocol. BetterToken’s Anthropic-compatible Base URL is
https://bettertoken.ai, and its request fields differ. Do not paste the OpenAI Python configuration into Claude Code. - Long-running agents: set a task budget, timeout, retry ceiling, and stop condition. Tool use does not justify unlimited filesystem or command permissions.
After configuring a tool, run three small checks: list the available models, send a short message, and complete a one-file repair with a test. Expand to repository-level tasks only after all three pass. This separates authentication, model-ID, and protocol problems from model-quality problems.
Run a reproducible test on your own work
Select 20 to 50 tasks with known correct outcomes. They should represent the work you actually expect the model to do, not a handful of examples selected after seeing which model completes them.
- Pin the repository commit, runtime, dependency cache, and tool permissions.
- Use the same prompt, timeout, and retry policy for every model.
- Evaluate
low,high, andmaxseparately instead of averaging them together. - Record input, cache hits and misses, visible output, retries, and total elapsed time.
- Count a patch as successful only after automated tests pass and a short human review accepts it.
| Field | Recommended record |
|---|---|
| Model ID | deepseek-flash |
reasoning_effort | low, high, or max |
max_tokens | Fixed by task class: 4K / 16K / 32K |
| Acceptance rule | Tests pass, no unrelated edits, requirements met |
| Token usage | input, cache hit/miss, output, retries |
| Timing | first response, green tests, human rework minutes |
Run at least two rounds so that cold caches, transient tool errors, or a brief service fluctuation do not dominate the result. Report success rate, accepted-task cost, and completion time together rather than selecting the single most flattering metric.
Choosing low, high, or max
low: everyday questions, explanations, small patches, and high-volume automation. It should normally be the default.high: difficult debugging, cross-file changes, and tasks that require more planning. Keep it only when the success-rate gain pays for the extra cost.max: the hardest agent tasks, architecture migrations, or infrequent high-value work. Apply a clear budget and timeout; do not make it the default for every request.- Fallback policy: start with
low; if it fails, retain the logs and test output and escalate tohigh; usemaxonly when the evidence points to insufficient reasoning rather than a broken environment or missing permission.
When a dependency cannot install, a test command is wrong, required files are absent, or the agent lacks write access, increasing the reasoning level usually raises cost without addressing the root cause.
Conclusion
The DeepSeek V4 Flash family is notable because it combines long context, low token pricing, and increasingly capable software-engineering performance in one product line. For current use, focus on V4.1 Flash and deepseek-flash; treat the V4 and 0731 architectures and scores as historical references.
The reliable decision order is simple: verify the release and parameters, compare official results under matching conditions, and then measure success rate, completion time, and cost per accepted task on your own repository. That produces a more useful answer than either “the cheapest price per million tokens” or “first place on one benchmark.”
Frequently asked questions
What is the context window of DeepSeek V4 Flash?
The official model cards for V4 Flash, 0731, and V4.1 Flash list a 1,000,000-token context window. A hosted endpoint may expose a lower limit, and input, history, tool results, and reserved output generally share that budget.
What should I set for DeepSeek V4 max_tokens?
Start around 4K for short coding tasks, 16K for longer changes, and consider 32K or more only for repository-scale work. DeepSeek’s current official API lists output of up to 384K, while the V4.1 model card recommends max_tokens >= 256K for local benchmark reproduction. Neither figure is a universal cap for every gateway or account.
Which model ID should I use now?
DeepSeek’s current API uses deepseek-flash. A legacy alias may temporarily route to V4.1, but production configurations should use the current ID and record the provider and date.
Do official benchmarks predict performance in Cursor or Aider?
Not directly. IDEs and agents also depend on prompts, tool implementation, repository structure, network access, permissions, test duration, and retry policy. Re-evaluate with your own task set.
Is DeepSeek V4.1 Flash good for coding?
Its official Terminal-Bench 2.1, DeepSWE v1.1, and NL2Repo-Bench results indicate strong software-engineering ability, and it offers a 1M context window plus reasoning controls. Suitability for a particular project still depends on observed success rate, latency, cost, and code review.
Sources
- DeepSeek V4 Flash official model card
- DeepSeek V4 Flash 0731 official model card
- DeepSeek V4.1 Flash official model card
- DeepSeek V4.1 Flash official announcement
- DeepSeek thinking-mode documentation
- DeepSeek official pricing
- Artificial Analysis: DeepSeek V4.1 Flash
- BetterToken API documentation
- BetterToken pricing