Use a Local Model in Claude Code with Ollama, Then Switch Back to Cloud

A practical guide for experienced Claude Code users: decide whether local inference fits the task, connect Qwen3.5 through Ollama's Anthropic-compatible API, validate file editing and command execution with a reversible one-file test, inspect context and CPU/GPU placement, understand compatibility gaps, and explicitly return to a cloud API when needed.

Contents
Use a Local Model in Claude Code with Ollama, Then Switch Back to Cloud

Claude Code can use a local model through Ollama’s Anthropic-compatible API, but a successful chat response does not prove that the model is ready for agentic coding. For a serious workflow, three things matter first: the model must support tool calls, the machine must sustain at least a 64k context window, and the task must be narrow enough to verify with commands and a diff.

This guide keeps the change reversible. You will connect Claude Code to qwen3.5 through Ollama’s official path, run a one-file acceptance task, check whether inference is actually local, review the API and data boundaries, and then remove the local override before returning to a cloud endpoint. The commands below use Bash on macOS, Linux, or WSL. They are procedures for you to run, not claims that this article executed the test on your hardware.

Decide first: local, cloud, or a hybrid workflow

Local models work best when the task is bounded and the result is mechanically verifiable. A large repository, cross-service migration, or difficult debugging session usually benefits more from a cloud model than from forcing an undersized local model to operate with heavy CPU offload.

WorkloadRecommended starting pointWhy
One-file fix, one new test, or explanation of a local functionTry local firstThe context is bounded and the result can be checked with a command and diff
Small or medium module with clear dependenciesLocal or hybridPass the smoke test, then expand the scope gradually
Large monorepo, cross-service refactor, or complex investigationCloud firstThese tasks need more effective context and more reliable tool planning
The model cannot hold 64k context without major CPU offloadCloud firstLatency and stalled interactions erase much of the local advantage
The workflow requires prompt caching, the Batches API, PDF blocks, or exact token countingCloud firstOllama currently implements only part of the Anthropic Messages API
Source code must not be sent to a remote modelLocal first, with cloud features disabledYou still need to audit web tools, MCP servers, and shell commands separately

A practical hybrid policy is to keep contained, repeatable edits local and switch explicitly to cloud for repository-wide reasoning, unsupported API features, or repeated local failures. That preserves one Claude Code interface without pretending that both backends behave identically.

Step 1: choose a tool-capable model and allocate 64k context

Claude Code needs more than text generation. The model must reliably produce tool calls so the client can read files, apply edits, and run commands. Ollama’s Qwen3.5 model page advertises tools support and includes a Claude Code launch command. You can also inspect the exact model you pulled through Ollama’s model-details API.

Pull the model and examine its capabilities:

ollama pull qwen3.5

curl http://localhost:11434/api/show \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3.5"}'

Before continuing, confirm that capabilities contains tools. If it does not, do not treat a normal chat response as a substitute for agent validation. Select a model that the current Ollama library explicitly marks as tool-capable, pull it, and repeat the check.

Context is the second gate. Ollama’s context-length documentation says that web search, agents, and coding tools should use at least 64,000 tokens, and that a larger context consumes more memory. In the Ollama app, set the context-length slider to 64000 or higher. For a service started from a shell, stop the existing instance first and run this in a dedicated terminal:

OLLAMA_CONTEXT_LENGTH=64000 ollama serve

Keep that terminal open. Wait for the server to start, then continue in a second terminal. If the port is already in use, an Ollama instance is already running; change the context setting for that instance instead of starting a second server.

Step 2: launch Claude Code through Ollama’s official integration

The shortest official route is:

ollama launch claude --model qwen3.5

This is the easiest way to establish the integration. After Claude Code starts, run /status and note the active settings sources. That record becomes useful later if a persistent settings layer keeps redirecting Claude Code to Ollama after you think you have switched back to cloud.

For a change that lasts only in the current terminal, configure the variables manually. The following remains Bash:

read -rs ANTHROPIC_AUTH_TOKEN
export ANTHROPIC_AUTH_TOKEN
export ANTHROPIC_API_KEY=""
export ANTHROPIC_BASE_URL="http://localhost:11434"
claude --model qwen3.5

read -rs accepts input without displaying it. Type ollama and press Enter. Ollama requires the authentication variable to be present but ignores its value on the local server. ANTHROPIC_BASE_URL routes model requests to the local Ollama endpoint, while --model qwen3.5 makes the test model explicit instead of leaving the result dependent on a stale ANTHROPIC_MODEL or a saved default.

Step 3: validate reading, editing, and command execution with one file

Do not make a production repository your first local-model test. Create an isolated directory where every result is visible through a file, exit status, and diff.

mkdir -p claude-ollama-smoke
cd claude-ollama-smoke
git init
cat > total.py <<'PY'
def total(values):
    return sum(values)

if __name__ == "__main__":
    assert total([2, 3]) == 5
PY
git add total.py
python3 total.py

python3 total.py should print nothing and exit with status 0. Start the local Claude Code session from this directory and send this task:

Modify only total.py.
If any item in values is not an int or float, make total raise TypeError with the exact message numbers only.
In __main__, add a check for [2, "3"] that confirms the same TypeError and message.
Run python3 total.py.
Do not modify any other file. Show the diff when finished.

The task is deliberately small, but it exercises the critical agent loop: read the file, plan an edit, invoke an editing tool, request a Bash command, observe the result, and present the final change. Keep Claude Code’s permission prompts enabled. A local model does not make unrestricted shell execution safe.

After the task, run these checks yourself:

python3 total.py
git status --short
git diff -- total.py
ollama ps

Use these acceptance criteria:

  1. python3 total.py exits with status 0.
  2. git status --short names only total.py, and git diff -- total.py contains only the requested type check and assertion.
  3. The Claude Code transcript shows file and Bash tool calls or permission prompts, rather than only a prose code suggestion.
  4. While the task is active, ollama ps lists qwen3.5, CONTEXT is at least 64000, and PROCESSOR shows whether the model is fully on GPU, partially offloaded, or primarily on CPU.

If any item fails, do not expand the scope to a real repository. Troubleshoot first, then decide whether to change models, narrow the task, or move to cloud.

Step 4: verify the execution boundary, not just the localhost URL

ANTHROPIC_BASE_URL=http://localhost:11434 shows that Claude Code sends model requests to a local port, but it does not prove that every part of the workflow is offline. Stronger evidence is the combination of a model tag without :cloud, the model appearing in ollama ps during the task, and local PROCESSOR and CONTEXT values that match the allocation on your machine.

Ollama’s FAQ states that it does not see prompts or data when a model runs locally, while prompts and responses for cloud-hosted models are processed by the cloud service. The current Qwen 3.5 page launches Claude Code with the local tag qwen3.5. Do not derive a cloud model name by appending a suffix to a local tag; when testing the cloud boundary, use a valid tag explicitly listed in the current official Cloud catalog or integration guide, such as gemma4:cloud. Determine where execution occurs from a valid model tag, ollama ps, and local resource allocation.

Audit the other network paths separately:

  • Commands invoked through Bash can access the network, upload files, or call another CLI.
  • MCP servers have their own processes, permissions, and data routes.
  • Web search, web fetch, and Ollama cloud models are not local inference.
  • Repository hooks, test scripts, and package managers may contact external services.

For a stricter local-only Ollama configuration, merge the following key into ~/.ollama/server.json without deleting other settings:

{
  "disable_ollama_cloud": true
}

Restart Ollama and check its logs for Ollama cloud disabled: true. Ollama documents that this disables its cloud models and web search. It still does not audit other network access performed by Claude Code, MCP servers, or shell commands.

Step 5: understand what the compatibility layer does not guarantee

Ollama exposes an Anthropic Messages API compatibility layer; it is not a complete reimplementation of the Anthropic API. The current documentation lists messages, streaming, system prompts, images, tool calls, tool results, and thinking among the supported capabilities, which is enough to build the basic Claude Code agent loop.

Protocol support is not behavioral parity. Tool-selection quality, patch quality, long-task stability, and respect for instructions still depend on the model, quantization, context allocation, and hardware. Passing the one-file test shows that the minimum path works in your environment; it does not prove that a local model will match a cloud Claude model on a large repository.

Ollama currently lists /v1/messages/count_tokens, prompt caching, the Batches API, citations, PDF document blocks, and server-sent streaming errors as unsupported. It also describes token counts as approximations based on the underlying tokenizer. If your workflow relies on one of those features, keep a cloud route available instead of discovering the gap halfway through a task.

Step 6: switch back to cloud explicitly

If the local variables exist only in the current Bash session, exit Claude Code and run:

unset ANTHROPIC_BASE_URL ANTHROPIC_AUTH_TOKEN ANTHROPIC_API_KEY ANTHROPIC_MODEL ANTHROPIC_DEFAULT_HAIKU_MODEL ANTHROPIC_DEFAULT_SONNET_MODEL ANTHROPIC_DEFAULT_OPUS_MODEL
claude

The new process can now follow your normal account login or cloud-provider configuration. Run /status after launch and ask one small, read-only question. A client that merely starts successfully is not proof that a cloud request has been completed.

If Claude Code still reaches Ollama, the override is probably stored in settings rather than the current shell. Claude Code’s official environment-variable reference says that an env value in a settings file overrides the same variable inherited from the shell. Use /status to identify active sources, then remove the local ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_API_KEY, and model overrides from the applicable layer:

  • ~/.claude/settings.json
  • .claude/settings.json
  • .claude/settings.local.json
  • organization-managed settings

Fully quit and restart Claude Code after the edit. A managed value cannot be defeated by a lower settings layer; an administrator must change it.

When a local model is unsuitable but you still want an Anthropic-compatible cloud API in the same Claude Code client, follow the current BetterToken Claude Code guide. The current Base URL is https://bettertoken.ai: it has no www and no /v1. Copy the exact Model ID from model plaza first. The current manual setup uses ANTHROPIC_MODEL for the primary model and the three ANTHROPIC_DEFAULT_*_MODEL variables for the Haiku, Sonnet, and Opus aliases. For a controlled smoke test, you can point all four variables to the same exact ID. This temporary Bash setup keeps the API key out of command history:

read -rsp "BetterToken API Key: " ANTHROPIC_AUTH_TOKEN
export ANTHROPIC_AUTH_TOKEN
read -rp $'\nBetterToken Model ID: ' ANTHROPIC_MODEL
export ANTHROPIC_MODEL
export ANTHROPIC_BASE_URL="https://bettertoken.ai"
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC="1"
export API_TIMEOUT_MS="3000000"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="$ANTHROPIC_MODEL"
export ANTHROPIC_DEFAULT_SONNET_MODEL="$ANTHROPIC_MODEL"
export ANTHROPIC_DEFAULT_OPUS_MODEL="$ANTHROPIC_MODEL"
claude

The first prompt hides the API-key input; at the second prompt, paste the exact Model ID copied from model plaza. This smoke test points the primary model and all three aliases to the same ID. If you deliberately use different models by role, set each default variable to its own exact ID instead. Do not append /v1 to the Base URL. After changing persistent settings, fully quit and restart Claude Code; for a temporary session, close the old process before running this block. Finally send one short, read-only request. Count the switch as complete only when it returns normally without a 401, connection, or model error and /status shows the expected active source; this does not imply full feature parity between local and cloud paths.

Troubleshoot the most common failure branches

ConnectionRefused or no response from localhost:11434

Confirm that the Ollama process is running and that the endpoint uses the expected port. Start it with ollama serve when needed. If the port is occupied, locate the existing instance instead of starting another one. Before reopening Claude Code, verify that curl http://localhost:11434/api/ps returns JSON.

Chat works, but Claude Code never reads or edits files

Call /api/show again and confirm that the model advertises tools. Then look for Claude Code permission prompts. If the model only writes “you could change the code like this” and never emits a tool call, switch to a model explicitly marked as tool-capable. A supported tool field proves transport compatibility, not reliable tool planning by every model.

The session is extremely slow or loses context on longer work

Run ollama ps and inspect PROCESSOR and CONTEXT. Heavy CPU offload, context below 64k, or repeated memory pressure are reasons to narrow the task, select a smaller tool-capable model, or use cloud. Do not remove permission and verification steps merely to make the interaction appear faster.

Shell changes do not change the endpoint or model

Run /status in Claude Code. A settings-file env value can replace the shell value, while --model and /model take precedence over ANTHROPIC_MODEL. Clean up the source that is actually winning, restart completely, and test with a read-only request.

The practical decision rule

Treat local Claude Code as an execution path that must earn a larger scope, not as a single toggle. Confirm tools, allocate at least 64k context, and use the one-file task to inspect tool calls, exit status, diff, and ollama ps. Expand only after those signals are stable.

When the task exceeds the machine, depends on an unsupported Anthropic feature, or repeatedly defeats the local model, remove the local endpoint and switch to cloud deliberately. A reliable rollback path is more valuable than forcing every coding task to remain local.

Ready to optimize your LLM workflow?

Join thousands of developers building faster, smarter, and more cost-effective AI applications with BetterToken.

Get Started for Free