Invite & Earn

How invite rewards work

Share your invite link. When a friend registers through it and tops up, you receive the displayed reward on their subsequent top-ups.

Hermes Agent Local Model Setup: Ollama, Qwen, and Explicit Cloud Switching

A practical, privacy-first Hermes Agent guide: connect a local Qwen model through Ollama, verify the setup with a small file task, configure a cloud provider, switch per session or for one turn, and audit every route that can send data off the machine.

Contents
Hermes Agent Local Model Setup: Ollama, Qwen, and Explicit Cloud Switching

You want Hermes Agent to handle everyday files and commands on your own machine, while keeping a cloud model available for the jobs that genuinely need it. The safest pattern is not an invisible automatic fallback: make local inference the default, prove that path with a small task, and turn every cloud escalation into a deliberate action.

This guide follows the current Hermes and Ollama documentation. It connects Hermes to a local Qwen model, runs a verifiable file task, configures a cloud provider, explains three switch scopes, and shows which model or tool steps can leave the machine. The commands are documentation-verified; they are not a claim that this exact hardware setup was benchmarked here.

Use a local model for routine work and a cloud model only after you decide the extra capability is worth the data boundary change.

SituationPreferWhy
Local file work, small code edits, routine tasks that can waitLocal Ollama/QwenModel requests go to 127.0.0.1, so the inference boundary is easy to inspect
Tool calls fail repeatedlyDiagnose the model and runtime firstA malformed tool call can be a server, parser, or context problem—not proof that the task is too hard
Long context, difficult reasoning, vision, or urgent latencyStart a clean session and switch to cloud explicitlyYou get stronger capacity without uploading an unrelated private session by accident
Nothing may leave the deviceLocal model plus local-only tools, no cloud auxiliaries, no fallbackA local main model does not make every tool or side task offline

Map the data boundary before changing any settings

Treat the main model, auxiliary models, and tools as separate routes. Any one of them can make an outbound request.

StepWhere data normally goesWhat to inspect
Hermes calls http://127.0.0.1:11434/v1Local OllamaThe model tag must not end in :cloud, and the endpoint must be loopback
The next turn after switching to a cloud modelThe selected cloud providerConversation context, system instructions, and required tool schemas may be included
Local file and terminal toolsUsually stays on the machineThe command itself may still upload, download, or call a remote API
Web, browser, search, remote MCP, or messaging toolsThe contacted serviceQueries, page content, tool arguments, and returned data
Local Whisper plus cloud TTSSpeech recognition is local; text for synthesis goes to cloudA mixed voice stack is not fully offline
fallback_providers or cloud auxiliary slotsCloud when triggeredErrors, capacity, auth, or routing conditions can change the boundary automatically

The Ollama Hermes integration lists both local and cloud models. A model tag ending in :cloud uses cloud inference even though you selected it through Ollama.

Fast path: let ollama launch hermes wire the endpoint

The current official Ollama integration provides a guided path:

ollama launch hermes

It can install Hermes when needed, offer a local-or-cloud model selector, and point Hermes at http://127.0.0.1:11434/v1. Choose an explicitly local entry. For example, qwen3.5:cloud is a cloud model because of the suffix.

After onboarding, verify the effective configuration rather than trusting the selector label:

hermes config get model --json
hermes status

Use the manual path below when the launcher cannot see your downloaded model or when you want to inspect every step.

Manual path: connect a local Qwen model step by step

1. Pull a Qwen model that advertises tool support

Ollama currently marks the Qwen 3.5 family as supporting tools. The example uses qwen3.5:4b because it is relatively small and easy to identify; it is a connectivity example, not a promise that a 4B model will be reliable for every agentic task. Pick a larger official tag when your RAM or VRAM allows it.

ollama --version
ollama pull qwen3.5:4b

Start the server only if your Ollama desktop service is not already running:

ollama serve

In another terminal, confirm that the local service can list models:

curl http://127.0.0.1:11434/api/tags

Test the OpenAI-compatible endpoint directly before involving Hermes:

curl http://127.0.0.1:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.5:4b",
    "messages": [{"role": "user", "content": "Reply with LOCAL_OK"}],
    "max_tokens": 16
  }'

A JSON response containing the model reply proves that Ollama is serving the local model. If this request is refused, fix the Ollama service first; changing Hermes will not solve it.

2. Point Hermes at the local endpoint

Run the full model wizard:

hermes model

Choose a custom endpoint and enter:

  • API Base URL: http://127.0.0.1:11434/v1
  • API Key: leave it blank; local Ollama does not require one
  • Model: qwen3.5:4b
  • Context length: leave it blank when your current Ollama integration offers auto-detection

Verify the saved route:

hermes config get model --json
hermes status

The result should identify the local model, the custom provider, and the loopback URL. Hermes documentation recommends a large context for tool-heavy agent work. If Hermes reports a context error, follow the current provider guidance for your version and inspect the value actually loaded by Ollama:

ollama ps

Do not force an oversized context if it causes out-of-memory errors or constant swapping.

Run a small file task and verify the result

Use a disposable directory so that success is observable and the agent does not need the web, messaging, or another remote service.

mkdir -p ~/hermes-local-check
cd ~/hermes-local-check
printf 'alpha\nbeta\ngamma\n' > input.txt
hermes

Inside the Hermes session, enter:

Read input.txt. Create summary.md with the line count and a one-sentence summary. Do not use web, browser, network, messaging, or cloud tools. Report the exact path you wrote.

After the agent finishes, inspect the output from another terminal or after exiting:

cd ~/hermes-local-check
cat summary.md

A useful pass condition has four parts:

  1. summary.md exists and recognizes three input lines.
  2. Hermes performed the file action instead of merely printing tool-call JSON.
  3. hermes status still reports local Qwen through 127.0.0.1.
  4. You did not configure an unwanted cloud fallback or cloud auxiliary model.

This proves that the local endpoint, tool loop, and file write are connected. It does not benchmark difficult tasks, and it does not prove that every optional tool is offline.

Configure a cloud provider without making cloud routing invisible

Run the model wizard again from a terminal:

hermes model

Choose the cloud provider you intend to use, complete its OAuth flow or interactive secret prompt, and select a model. Do not paste a key into a chat message or a shell command that will be saved in history.

Because hermes model changes the default, switch the active session and future default back to local after the cloud provider is configured:

/model qwen3.5:4b --provider custom --global

You now have three explicit cloud switch scopes:

/model YOUR_CLOUD_MODEL --provider YOUR_CLOUD_PROVIDER
/model YOUR_CLOUD_MODEL --provider YOUR_CLOUD_PROVIDER --once
/model YOUR_CLOUD_MODEL --provider YOUR_CLOUD_PROVIDER --global
  • The first command changes only the current session.
  • --once uses the cloud model for the next turn and then restores the previous model.
  • --global changes the running session and the default for new sessions.

Return the current session to local with:

/model qwen3.5:4b --provider custom

If a provider is missing from /model, configure or authenticate it with hermes model. The in-session command switches among configured choices; it does not create a new credentialed provider.

Decide when a cloud switch is justified

Stay local for private, bounded work

Keep local inference for unreleased source code, personal documents, internal logs, and small edits with a clear finish line. If the model calls the required tools correctly and the speed is acceptable, a theoretical quality gain is not a good reason to upload the session.

Switch to cloud after ruling out setup problems

A cloud model is reasonable when the local model still fails after you have confirmed the endpoint, context, and tool support; when the task needs much longer context, stronger multi-step reasoning, or vision; or when local latency is unacceptable.

For sensitive work, start a fresh cloud session and provide only the minimum necessary material. Hermes documents that a mid-session model change resets the prompt cache and the next request re-reads the conversation. From a privacy perspective, assume that the relevant current-session content is sent to the cloud provider after the switch.

Audit auxiliary models and fallbacks, not just the main model

If your policy is “every outbound model call requires my decision,” do not configure fallback_providers. A fallback can activate after errors, rate limits, authentication failures, or capacity problems; it is not the same as a human deciding that a task is difficult enough for cloud use.

Hermes can also route compression, vision, web extraction, and other side jobs to auxiliary models. A local main model does not prove that those slots are local.

On Bash, macOS, or Linux, this read-only check finds the relevant routing fields without intentionally printing an API-key field:

grep -nE 'fallback|auxiliary|base_url|provider|default' ~/.hermes/config.yaml

On Windows, inspect %USERPROFILE%\.hermes\config.yaml in a text editor. Check the main endpoint, auxiliary providers, fallback_providers, and any remote Base URL you do not recognize.

For stronger evidence, use host firewall logs, DNS logs, or an outbound proxy to observe actual connections. A UI badge that says “local” is not a packet-level guarantee for the whole workflow.

Troubleshooting the common failure modes

Connection refused

Test Ollama directly:

curl http://127.0.0.1:11434/api/tags

If it fails, start ollama serve or the Ollama desktop service before touching Hermes settings.

The model prints tool-call JSON instead of acting

Confirm that the exact Qwen tag advertises tools and that Hermes and Ollama are current. Small models can still be unreliable on complex arguments even when the family supports tools. Re-run the one-file test, then try a larger local model or make a deliberate cloud switch.

The model or provider is absent

Use hermes model in the terminal to add the custom endpoint or authenticate the cloud provider. /model only exposes choices Hermes already knows.

The current chat still behaves like the old model

Changing the dashboard default or running hermes model mainly affects new sessions. Hot-swap an active chat with /model, or start a clean session.

Responses are slow or context errors appear

Inspect the loaded model:

ollama ps

Check GPU offload, model size, and context. A larger context can help an agent retain tool schemas, but it also increases memory use and first-turn prefill time.

A hybrid stack can be useful without being fully offline

One X user described a personal prototype combining Gemma through Ollama Cloud, cloud Qwen, local Qwen, local Whisper, and cloud text-to-speech. That is a useful illustration of per-task routing, but it is one person’s report, without a published reproducible configuration, network log, or independent performance test.

More importantly, models, speech, and tools have separate boundaries. Local speech recognition does not make cloud TTS local, and a local main model does not make calendar, project-management, search, or remote MCP calls private by default.

Final checklist

Before relying on the setup, confirm all of the following:

  • The local model tag does not end in :cloud, and the Base URL is http://127.0.0.1:11434/v1.
  • The direct curl test succeeds, and hermes status shows the expected route.
  • The disposable file task actually creates summary.md.
  • Cloud providers are configured, but every switch is explicit through /model.
  • Sensitive cloud work starts in a fresh session with reduced context.
  • Remote tools, auxiliary models, and fallback_providers have been reviewed.
  • When the requirement is fully offline, all network tools and cloud services remain disabled.

Official documentation and the user-report source

Ready to optimize your LLM workflow?

Join thousands of developers building faster, smarter, and more cost-effective AI applications with BetterToken.

Get Started for Free