Qwen Code Usage Limit Reached? Identify the Quota and Keep Working
Qwen Code does not have one universal quota. This guide shows how to identify the active authentication source, check the matching limit and reset rule, distinguish a real quota stop from configuration errors, and choose a supported way to continue.
Contents

When Qwen Code says that usage is exhausted, do not assume there is one universal Qwen Code quota. Qwen Code is the client. The limit comes from the authentication route and provider currently serving the request.
Start inside Qwen Code with /doctor to identify the current authentication. Then check the configured key name and baseUrl without displaying the secret itself. Only after you know whether the session uses Qwen OAuth, Alibaba Cloud Coding Plan, Token Plan, a standard API key, or a third-party/custom provider can you check the right quota and recovery rule. The old qwen auth status workflow has been removed; current official guidance points to /doctor and /auth.
Also keep this separate from Qwen3.8 Max API pricing. A model price tells you how token billing works for one API route. It does not tell you which quota Qwen Code is currently consuming.
First, identify which quota you are actually using
| Authentication source | Clues in the configuration | What “exhausted” usually refers to | Supported way to continue |
|---|---|---|---|
| Qwen OAuth | Old browser-login credentials; no current selectable OAuth entry in /auth | The former free tier was discontinued, not temporarily depleted | Run /auth and select a current provider |
| Alibaba Cloud Coding Plan | Plan-specific key plus a coding...dashscope.aliyuncs.com endpoint | One of the plan’s request caps, or a separate technical rate/capacity limit | Wait for the applicable cap to replenish, or switch authentication and accept separate billing |
| Alibaba Cloud Token Plan | Token Plan key plus a dedicated token-plan...maas.aliyuncs.com endpoint | Credits or plan allowance for the active edition, seat, or shared pack | Wait for the next cycle, add permitted quota, or switch authentication |
| Standard Model Studio API key | General regional Model Studio endpoint rather than a Coding Plan endpoint | Provider balance, model-specific free quota, token billing, or rate limit | Enable or fund billing, wait for a rate limit, or select another supported provider |
| Third-party or custom provider | Provider-specific baseUrl, protocol, model and environment variable | That provider’s credits, subscription, account limit, or rate limit | Follow that provider’s dashboard and billing rules, or change provider with /auth |
The model name alone is not enough. The same model can be exposed through different plans and endpoints, and those routes do not share quota or billing.
A safe five-minute diagnosis
1. Record the exact failure before changing anything
Save the error text, time, selected model and whether the failure appeared immediately or after a long task. Do not paste an API key into a ticket or screenshot. A message that explicitly says a plan quota is exhausted is stronger evidence than a generic 401, 403, 429, timeout or “model not found” response.
2. Run /doctor inside Qwen Code
The official authentication guide names /doctor as the current way to check authentication. Record the authentication type, provider and endpoint category that it reports. If the session is still tied to the discontinued Qwen OAuth route, there is no free-tier reset to wait for.
Use /auth when you need to change authentication. Use /model only to choose among models already configured for a provider. Changing a model does not automatically move the session to a different billing account or plan.
3. Inspect both settings scopes
Check the user file ~/.qwen/settings.json and the project file .qwen/settings.json. Project settings can override user settings. Environment variables and command-line arguments have still higher precedence, so a stale variable can keep routing Qwen Code to an old provider even after you edit the JSON file.
Look only for these non-secret facts:
- the selected authentication protocol;
- the environment-variable name, not its value;
- the provider
baseUrl; - the selected model ID;
- any Coding Plan or Token Plan region marker.
If a key is stored directly in settings.json, move it to a protected secret source where practical. Never print the value just to prove that it exists.
4. Match the endpoint to the plan
For the international Coding Plan, the documented OpenAI-compatible endpoint contains coding-intl.dashscope.aliyuncs.com. The international Token Plan uses a dedicated token-plan.ap-southeast-1.maas.aliyuncs.com endpoint. Standard pay-as-you-go Model Studio uses regional endpoints instead.
A Coding Plan key and a general Model Studio key are not interchangeable. Mixing a general key with a pay-as-you-go endpoint can create separate charges rather than consuming the Coding Plan quota. Mixing the wrong key and endpoint can also look like an authentication problem instead of a quota problem.
5. Open the matching provider console
Do not infer remaining quota from the Qwen Code terminal alone. Open the console for the account, region and plan identified above. Check remaining usage, the cap that was reached, reset time, subscription status, billing status and any provider incident or throttling notice.
If you use Alibaba Cloud Coding Plan
As of October 1, 2026, the official international Coding Plan page lists the Pro plan at USD 50 per month with three simultaneous caps:
- 6,000 requests per five hours;
- 45,000 requests per week;
- 90,000 requests per month.
Whichever cap is reached first stops further calls. A user instruction is not necessarily one request: the official page explains that one query can trigger multiple model calls, and complex agent work can consume much more quota than a short question.
The reset rules are different for each cap:
- the five-hour allowance is rolling, so usage is released when it becomes five hours old;
- the weekly allowance resets every Monday at 00:00 (UTC+08:00);
- the monthly allowance resets on the subscription renewal date at 00:00 (UTC+08:00).
This distinction matters. Waiting five hours will not repair a weekly or monthly cap. Check the Coding Plan console to see which cap is active. Accounts in other regions can have different purchase pages, currencies or availability, so use the page and console tied to the account’s region.
To continue before a reset, you may change to a standard API key, Token Plan, or another provider through /auth. That creates a separate quota and billing path; it does not extend the Coding Plan. Keep the new provider’s key and endpoint together.
Coding Plan is intended for interactive use in supported coding tools. The official terms on the plan page prohibit using its key for automated scripts or application backends.
If you use Token Plan
Token Plan is not the Coding Plan request counter. Current international Alibaba Cloud documentation describes Token Plan in Credits, with editions and allowances that can vary by account and product version. Credits consumed depend on the model, token count, thinking mode and tool calls.
For the documented Team Edition, usage draws from the assigned seat first and then from an available shared quota pack. When all Credits are exhausted, service is suspended until the next billing cycle or until an allowed quota pack is added. The console shows usage percentage, reset time, seat allocation and pack status.
Therefore:
- confirm the Token Plan edition and region in the console;
- check the active seat or account Credits rather than Coding Plan request caps;
- confirm that Qwen Code is using the dedicated Token Plan endpoint;
- add quota only through the official console if the edition permits it, or wait for the displayed cycle;
- use
/authif you intentionally want to move to a different billing source.
Do not copy a Coding Plan reset schedule onto Token Plan. They are different products.
If you use a standard API key
A standard Model Studio API key is a pay-as-you-go route. Its practical limit comes from the account’s balance or billing status, model- and region-specific free quota, provider rate limits and any service capacity controls. There is no universal “Qwen Code weekly reset” for this route.
Check the Model Studio billing and usage pages for the exact account and region. If the account has no billable balance or an applicable free quota has ended, enable or fund billing according to the provider’s current rules. If the response is a rate limit, wait for the documented window instead of changing keys repeatedly.
A 401 or 403 usually points first to authentication, permissions, region or endpoint mismatch. A 429 can mean a quota or rate limit, but it can also reflect temporary throttling. Read the response body and provider dashboard before classifying it.
If you use a third-party or custom provider
Qwen Code supports built-in third-party providers and custom OpenAI-, Anthropic- and Gemini-compatible endpoints. In this setup, Qwen Code does not define the commercial quota. The provider does.
Open that provider’s dashboard and verify:
- account balance or subscription allowance;
- the model’s availability under the account;
- per-minute or concurrent-request limits;
- the reset or renewal time;
- the exact endpoint and model ID expected by the provider.
Switching to another model on the same provider may still use the same shared balance. Switching providers with /auth is a real billing change and should be deliberate.
Quota exhaustion or a configuration error?
Use the following order because it avoids unnecessary purchases and reconfiguration:
- Explicit OAuth-discontinued message: do not wait; choose a current authentication method.
- Explicit plan-cap message plus a matching console cap: follow that cap’s reset or quota-purchase rule.
401or403: verify key, permissions, endpoint and region before treating it as exhaustion.- “Model not found” or unsupported model: verify the model allowlist and endpoint; this is not evidence of an empty quota.
429, queueing or temporary interruption: check rate limits,Retry-After, status notices and the console. A plan can have technical concurrency or rate controls in addition to its headline quota.- Timeout or network error: test connectivity and provider status; quota may be unrelated.
How to verify that the recovery worked
After changing authentication or waiting for a reset:
- restart Qwen Code so changed environment variables are reloaded;
- run
/doctorand confirm the intended authentication source; - run
/modeland select a model actually supported by that provider; - send a small, read-only request, such as asking Qwen Code to summarize one file without editing it;
- confirm success in the terminal and, where available, verify that usage appeared in the expected provider console;
- only then resume a long agent task.
A successful small request proves that authentication, endpoint and model selection work together. It does not guarantee that a large task will fit inside the remaining quota.
Common mistakes
- Waiting five hours when the weekly or monthly Coding Plan cap was reached.
- Assuming a restart resets a provider quota.
- Changing only the key or only the endpoint and leaving an invalid pair.
- Editing user settings while a project setting or environment variable still overrides them.
- Using
/modelas if it were a billing switch. - Treating Qwen3.8 Max token pricing as the Qwen Code plan limit.
- Repeatedly rotating keys when the provider account itself has no balance or allowance.
- Using a Coding Plan key in automated scripts, which is outside the documented usage scope.
Frequently asked questions
Is there still a universal free Qwen Code quota?
No. The old Qwen OAuth free tier was discontinued on April 15, 2026. A standard API provider may offer its own temporary free quota or credits, but those belong to that provider, model, region and account.
Will restarting Qwen Code restore usage?
No. Restarting reloads configuration; it does not replenish a provider quota. It helps only after credentials or settings changed, or after the provider has already reset the allowance.
Can I bypass the limit by switching models?
Only when the new model is charged to a genuinely different allowance. Models under the same plan can share a cap or Credits pool. Confirm in the provider console.
Should I create a new API key?
Only if the current key is invalid, revoked or exposed. A new key under the same exhausted account usually does not create new quota.
Where should I check first next time?
Run /doctor, identify the endpoint and provider, then open that provider’s usage page. This sequence is faster and safer than guessing from the model name.