Hermes Agent Local Model Setup: Ollama, Qwen, and Explicit Cloud Switching
A practical, privacy-first Hermes Agent guide: connect a local Qwen model through Ollama, verify the setup with a small file task, configure a cloud provider, switch per session or for one turn, and audit every route that can send data off the machine.
Contents

You want Hermes Agent to handle everyday files and commands on your own machine, while keeping a cloud model available for the jobs that genuinely need it. The safest pattern is not an invisible automatic fallback: make local inference the default, prove that path with a small task, and turn every cloud escalation into a deliberate action.
This guide follows the current Hermes and Ollama documentation. It connects Hermes to a local Qwen model, runs a verifiable file task, configures a cloud provider, explains three switch scopes, and shows which model or tool steps can leave the machine. The commands are documentation-verified; they are not a claim that this exact hardware setup was benchmarked here.
The recommended policy: local by default, cloud by explicit choice
Use a local model for routine work and a cloud model only after you decide the extra capability is worth the data boundary change.
| Situation | Prefer | Why |
|---|---|---|
| Local file work, small code edits, routine tasks that can wait | Local Ollama/Qwen | Model requests go to 127.0.0.1, so the inference boundary is easy to inspect |
| Tool calls fail repeatedly | Diagnose the model and runtime first | A malformed tool call can be a server, parser, or context problem—not proof that the task is too hard |
| Long context, difficult reasoning, vision, or urgent latency | Start a clean session and switch to cloud explicitly | You get stronger capacity without uploading an unrelated private session by accident |
| Nothing may leave the device | Local model plus local-only tools, no cloud auxiliaries, no fallback | A local main model does not make every tool or side task offline |
Map the data boundary before changing any settings
Treat the main model, auxiliary models, and tools as separate routes. Any one of them can make an outbound request.
| Step | Where data normally goes | What to inspect |
|---|---|---|
Hermes calls http://127.0.0.1:11434/v1 | Local Ollama | The model tag must not end in :cloud, and the endpoint must be loopback |
| The next turn after switching to a cloud model | The selected cloud provider | Conversation context, system instructions, and required tool schemas may be included |
| Local file and terminal tools | Usually stays on the machine | The command itself may still upload, download, or call a remote API |
| Web, browser, search, remote MCP, or messaging tools | The contacted service | Queries, page content, tool arguments, and returned data |
| Local Whisper plus cloud TTS | Speech recognition is local; text for synthesis goes to cloud | A mixed voice stack is not fully offline |
fallback_providers or cloud auxiliary slots | Cloud when triggered | Errors, capacity, auth, or routing conditions can change the boundary automatically |
The Ollama Hermes integration lists both local and cloud models. A model tag ending in :cloud uses cloud inference even though you selected it through Ollama.
Fast path: let ollama launch hermes wire the endpoint
The current official Ollama integration provides a guided path:
ollama launch hermes
It can install Hermes when needed, offer a local-or-cloud model selector, and point Hermes at http://127.0.0.1:11434/v1. Choose an explicitly local entry. For example, qwen3.5:cloud is a cloud model because of the suffix.
After onboarding, verify the effective configuration rather than trusting the selector label:
hermes config get model --json
hermes status
Use the manual path below when the launcher cannot see your downloaded model or when you want to inspect every step.
Manual path: connect a local Qwen model step by step
1. Pull a Qwen model that advertises tool support
Ollama currently marks the Qwen 3.5 family as supporting tools. The example uses qwen3.5:4b because it is relatively small and easy to identify; it is a connectivity example, not a promise that a 4B model will be reliable for every agentic task. Pick a larger official tag when your RAM or VRAM allows it.
ollama --version
ollama pull qwen3.5:4b
Start the server only if your Ollama desktop service is not already running:
ollama serve
In another terminal, confirm that the local service can list models:
curl http://127.0.0.1:11434/api/tags
Test the OpenAI-compatible endpoint directly before involving Hermes:
curl http://127.0.0.1:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.5:4b",
"messages": [{"role": "user", "content": "Reply with LOCAL_OK"}],
"max_tokens": 16
}'
A JSON response containing the model reply proves that Ollama is serving the local model. If this request is refused, fix the Ollama service first; changing Hermes will not solve it.
2. Point Hermes at the local endpoint
Run the full model wizard:
hermes model
Choose a custom endpoint and enter:
- API Base URL:
http://127.0.0.1:11434/v1 - API Key: leave it blank; local Ollama does not require one
- Model:
qwen3.5:4b - Context length: leave it blank when your current Ollama integration offers auto-detection
Verify the saved route:
hermes config get model --json
hermes status
The result should identify the local model, the custom provider, and the loopback URL. Hermes documentation recommends a large context for tool-heavy agent work. If Hermes reports a context error, follow the current provider guidance for your version and inspect the value actually loaded by Ollama:
ollama ps
Do not force an oversized context if it causes out-of-memory errors or constant swapping.
Run a small file task and verify the result
Use a disposable directory so that success is observable and the agent does not need the web, messaging, or another remote service.
mkdir -p ~/hermes-local-check
cd ~/hermes-local-check
printf 'alpha\nbeta\ngamma\n' > input.txt
hermes
Inside the Hermes session, enter:
Read input.txt. Create summary.md with the line count and a one-sentence summary. Do not use web, browser, network, messaging, or cloud tools. Report the exact path you wrote.
After the agent finishes, inspect the output from another terminal or after exiting:
cd ~/hermes-local-check
cat summary.md
A useful pass condition has four parts:
summary.mdexists and recognizes three input lines.- Hermes performed the file action instead of merely printing tool-call JSON.
hermes statusstill reports local Qwen through127.0.0.1.- You did not configure an unwanted cloud fallback or cloud auxiliary model.
This proves that the local endpoint, tool loop, and file write are connected. It does not benchmark difficult tasks, and it does not prove that every optional tool is offline.
Configure a cloud provider without making cloud routing invisible
Run the model wizard again from a terminal:
hermes model
Choose the cloud provider you intend to use, complete its OAuth flow or interactive secret prompt, and select a model. Do not paste a key into a chat message or a shell command that will be saved in history.
Because hermes model changes the default, switch the active session and future default back to local after the cloud provider is configured:
/model qwen3.5:4b --provider custom --global
You now have three explicit cloud switch scopes:
/model YOUR_CLOUD_MODEL --provider YOUR_CLOUD_PROVIDER
/model YOUR_CLOUD_MODEL --provider YOUR_CLOUD_PROVIDER --once
/model YOUR_CLOUD_MODEL --provider YOUR_CLOUD_PROVIDER --global
- The first command changes only the current session.
--onceuses the cloud model for the next turn and then restores the previous model.--globalchanges the running session and the default for new sessions.
Return the current session to local with:
/model qwen3.5:4b --provider custom
If a provider is missing from /model, configure or authenticate it with hermes model. The in-session command switches among configured choices; it does not create a new credentialed provider.
Decide when a cloud switch is justified
Stay local for private, bounded work
Keep local inference for unreleased source code, personal documents, internal logs, and small edits with a clear finish line. If the model calls the required tools correctly and the speed is acceptable, a theoretical quality gain is not a good reason to upload the session.
Switch to cloud after ruling out setup problems
A cloud model is reasonable when the local model still fails after you have confirmed the endpoint, context, and tool support; when the task needs much longer context, stronger multi-step reasoning, or vision; or when local latency is unacceptable.
For sensitive work, start a fresh cloud session and provide only the minimum necessary material. Hermes documents that a mid-session model change resets the prompt cache and the next request re-reads the conversation. From a privacy perspective, assume that the relevant current-session content is sent to the cloud provider after the switch.
Audit auxiliary models and fallbacks, not just the main model
If your policy is “every outbound model call requires my decision,” do not configure fallback_providers. A fallback can activate after errors, rate limits, authentication failures, or capacity problems; it is not the same as a human deciding that a task is difficult enough for cloud use.
Hermes can also route compression, vision, web extraction, and other side jobs to auxiliary models. A local main model does not prove that those slots are local.
On Bash, macOS, or Linux, this read-only check finds the relevant routing fields without intentionally printing an API-key field:
grep -nE 'fallback|auxiliary|base_url|provider|default' ~/.hermes/config.yaml
On Windows, inspect %USERPROFILE%\.hermes\config.yaml in a text editor. Check the main endpoint, auxiliary providers, fallback_providers, and any remote Base URL you do not recognize.
For stronger evidence, use host firewall logs, DNS logs, or an outbound proxy to observe actual connections. A UI badge that says “local” is not a packet-level guarantee for the whole workflow.
Troubleshooting the common failure modes
Connection refused
Test Ollama directly:
curl http://127.0.0.1:11434/api/tags
If it fails, start ollama serve or the Ollama desktop service before touching Hermes settings.
The model prints tool-call JSON instead of acting
Confirm that the exact Qwen tag advertises tools and that Hermes and Ollama are current. Small models can still be unreliable on complex arguments even when the family supports tools. Re-run the one-file test, then try a larger local model or make a deliberate cloud switch.
The model or provider is absent
Use hermes model in the terminal to add the custom endpoint or authenticate the cloud provider. /model only exposes choices Hermes already knows.
The current chat still behaves like the old model
Changing the dashboard default or running hermes model mainly affects new sessions. Hot-swap an active chat with /model, or start a clean session.
Responses are slow or context errors appear
Inspect the loaded model:
ollama ps
Check GPU offload, model size, and context. A larger context can help an agent retain tool schemas, but it also increases memory use and first-turn prefill time.
A hybrid stack can be useful without being fully offline
One X user described a personal prototype combining Gemma through Ollama Cloud, cloud Qwen, local Qwen, local Whisper, and cloud text-to-speech. That is a useful illustration of per-task routing, but it is one person’s report, without a published reproducible configuration, network log, or independent performance test.
More importantly, models, speech, and tools have separate boundaries. Local speech recognition does not make cloud TTS local, and a local main model does not make calendar, project-management, search, or remote MCP calls private by default.
Final checklist
Before relying on the setup, confirm all of the following:
- The local model tag does not end in
:cloud, and the Base URL ishttp://127.0.0.1:11434/v1. - The direct
curltest succeeds, andhermes statusshows the expected route. - The disposable file task actually creates
summary.md. - Cloud providers are configured, but every switch is explicit through
/model. - Sensitive cloud work starts in a fresh session with reduced context.
- Remote tools, auxiliary models, and
fallback_providershave been reviewed. - When the requirement is fully offline, all network tools and cloud services remain disabled.
Official documentation and the user-report source
- Hermes: local Ollama setup
- Hermes: configuring and switching models
- Hermes: providers and custom endpoints
- Ollama: Hermes Agent integration
- Ollama: Qwen 3.5 model library
- X post describing the hybrid prototype (a user report, not an independent benchmark)