Hermes reasoning effort: session, global, and per-model settings

A reproducible way to choose Hermes reasoning effort: separate the thinking display from the actual effort, configure session, global, and per-model scopes, then compare one fixed task using quality, latency, and provider-side usage.

Contents
Hermes reasoning effort: session, global, and per-model settings

Keeping Hermes at its highest reasoning level does not make every task better, and visible thinking is not proof that the request used the level you selected. A safer workflow is to keep a practical global default, raise effort temporarily for difficult work, add per-model defaults only after repeated evidence, and verify the result with one checkable task plus provider-side usage records.

The practical default: start at medium

Hermes currently accepts none, minimal, low, medium, high, xhigh, max, and ultra. An unset value resolves to medium. A model or route may support only part of that ladder, so the value can be clamped, translated, ignored, or rejected; the label alone is never a performance guarantee. Check the current Hermes configuration documentation and your provider’s request records.

TaskSensible starting pointWhen to move
Formatting, extraction, or a deterministic rewritelow; test minimal or none only after confirming supportMove to medium if fields or constraints are missed
A small code edit, routine question, or clearly scoped debuggingmediumTry low after consistently correct runs; try high after missed constraints
Multi-constraint review, cross-file diagnosis, or trade-off analysishighTest xhigh or max only if the quality gain repeats and the delay is acceptable
Exceptionally hard planning or long reasoning chainsCompare high with xhigh firstKeep max or ultra only when a controlled comparison shows a useful gain

ultra is an internal Hermes rung. The route maps it to the strongest value it can actually send, so it is a poor choice for a blind global default.

Two different controls: showing thinking is not reasoning effort

These commands change the current session’s reasoning effort:

/reasoning high
/reasoning none

These commands only change whether thinking is shown in the interface:

/reasoning show
/reasoning hide

A hidden trace can still be produced by a request running at high. A visible trace does not prove that the request used a high level. Enter /reasoning with no argument to inspect both the current effort and the display state instead of inferring one from the other.

When to use session, global, or per-model scope

Session scope: solve the task in front of you

Inside an active session, run:

/reasoning high

The change is session-scoped by default. This is the safest way to give one difficult debugging or architecture task more effort without changing every future conversation.

To request reasoning off for the current session, use:

/reasoning none

That only turns reasoning off if the selected model and route support it. A provider may require reasoning, translate the setting, or reject it, so the request record still matters.

Global scope: choose the everyday default

Add --global to persist a default for new sessions:

/reasoning medium --global

Hermes saves this as agent.reasoning_effort. For mixed workloads, medium is a safer global baseline than the maximum level; use session overrides for the minority of tasks that need more.

Read the saved configuration from the terminal:

hermes config path
hermes config get agent.reasoning_effort
hermes config check

A successful config get proves that Hermes resolved a configuration value. It does not, by itself, prove that the provider accepted or applied that value.

Per-model scope: stable defaults for models you switch between

If you regularly switch between a fast model and a deeper reasoning model, edit config.yaml:

agent:
  reasoning_effort: "medium"
  reasoning_overrides:
    "custom/example-fast-model": "low"
    "custom/example-deep-model": "high"

A matching per-model override takes priority over the global agent.reasoning_effort. Prefer the exact model ID configured in Hermes. After editing the file, start a fresh session, select the target model, and run /reasoning again.

You can inspect the map with:

hermes config get agent.reasoning_overrides --json

Model IDs often contain dots and slashes. Direct YAML editing is straightforward; when creating a dotted key through hermes config set, follow the literal-dot escaping rules in the CLI command reference.

Precedence: why changing the global value may appear to do nothing

For a selected model, think of the effective order as:

  1. the current session’s temporary /reasoning choice;
  2. a matching agent.reasoning_overrides entry;
  3. the global agent.reasoning_effort;
  4. the model or provider default.

If the global value is low but /reasoning still reports high, first look for a live session override or a per-model entry. Recheck after /model switches as well, because the new model can match a different override.

Use one fixed, checkable task

Do not test one level on a trivial rewrite and another on a hard bug. That measures task differences, not effort. The following small task has a result that can be checked by hand and does not need tools:

The function should merge overlapping or touching closed integer ranges without shrinking coverage.
Find one minimal counterexample, give the expected and actual output, make the smallest code fix, and add three regression tests.
Do not use tools. Return JSON only, with the keys counterexample, expected, actual, fix, and tests.

def merge_ranges(ranges):
    ranges = sorted(ranges)
    merged = []
    for start, end in ranges:
        if not merged or start > merged[-1][1] + 1:
            merged.append([start, end])
        else:
            merged[-1][1] = end
    return merged

The key failure occurs when a later range is fully contained in the current one: assigning the smaller end can shrink coverage. Score each run on five objective checks rather than style:

  1. valid JSON and no extra prose;
  2. a contained-range counterexample that really triggers the bug;
  3. correct expected and actual output;
  4. a minimal fix that preserves the larger end;
  5. regression tests for contained, touching, and disjoint ranges.

Manual comparison: the fastest way to choose a lasting default

Use a fresh session for each candidate level. Keep the model ID, provider, working directory, context, tool settings, task text, and output format unchanged. One run is enough for screening; if the result will change a frequent default, run each finalist at least three clean times so a random variation is not mistaken for a stable gain.

Record the following:

FieldHow to record it
EffortEnter /reasoning before the task and save the reported value
QualityUse the 0–5 checklist above
LatencyWall-clock time from submission to the final answer
Model and providerConfirm in Hermes status and the provider request record
Actual usageUse the provider/API request record or billing detail
AnomaliesMark timeouts, retries, fallbacks, errors, or model switches

Do not average a retry or fallback run with clean runs. It can change the model, call count, latency, and token usage at the same time, which makes it unsuitable for judging reasoning effort.

Save a local Hermes report with --usage-file

For machine-readable comparisons, temporarily change the global level and run the same one-shot task:

hermes config set agent.reasoning_effort low
hermes -z "Review the supplied merge_ranges function and return the requested JSON only." --usage-file ./hermes-low-usage.json > ./hermes-low-output.txt

hermes config set agent.reasoning_effort medium
hermes -z "Review the supplied merge_ranges function and return the requested JSON only." --usage-file ./hermes-medium-usage.json > ./hermes-medium-output.txt

hermes config set agent.reasoning_effort high
hermes -z "Review the supplied merge_ranges function and return the requested JSON only." --usage-file ./hermes-high-usage.json > ./hermes-high-output.txt

In a real comparison, pass the same complete task each time rather than the shortened prompt above. Before running, make sure the selected model has no per-model override that would shadow the global value. Restore the original setting afterward. If it was previously unset, run:

hermes config unset agent.reasoning_effort

Otherwise, set it back to the recorded original value.

The Hermes JSON report can include input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, reasoning_tokens, total_tokens, api_calls, model, provider, and estimated_cost_usd. Top-level counters cover the main agent loop. Auxiliary calls such as title generation, vision, or compression are separated under auxiliary; total_including_auxiliary is the local combined total.

Keep three boundaries clear:

  • estimated_cost_usd is a local estimate, not the provider’s invoice;
  • if a provider does not return a token category, a missing field is not evidence of zero usage;
  • if a run retries or falls back, confirm the real calls and models in the provider’s request list.

Separate four evidence types

A useful verification records four different things:

  1. Configuration readback: hermes config get and config.yaml contain the intended value. This proves what Hermes stored and resolved, not what the provider accepted.
  2. Effort actually sent or mapped: after selecting the target model, run /reasoning. Where the status shows sends ... on this route, or the route exposes an outbound request trace, verify that the value sent to the API matches the expected mapping. Thinking visibility remains a separate display setting.
  3. Provider-side receipt, acceptance, or execution: use a server-side request record, echoed value, or explicit acceptance/execution confirmation only when the route or provider actually exposes it. A success signal is that the server-side parameter matches the sent/mapped value and no rejection, retry, fallback, or further downgrade is recorded. A payload trace proves receipt, not execution; do not require fields the provider does not expose.
  4. Usage and outcome: record the actual model, output quality, latency, token categories, API calls, and billed or estimated cost. These can verify routing and real usage, but not which effort the provider accepted; the number of reasoning tokens cannot be reverse-mapped to a tier.

These evidence types do not replace one another. If provider logs show only the model, tokens, call count, or spend and do not expose effort, the defensible conclusion is that you compared output, latency, and real usage under the recorded configurations, but could not confirm the effort actually accepted by the provider.

Troubleshooting when none appears not to stick

A public GitHub issue filed on October 5, 2026 reported that, on the specific main commits named by the reporter, hermes config set agent.reasoning_effort none could store YAML null while /reasoning none --global stored the string none. This is a version-bounded user report. It does not prove that your current version still has the behavior, and it does not prove that a later version fixed it.

Use this check order:

  1. run /reasoning none --global;
  2. run hermes config get agent.reasoning_effort;
  3. open the file returned by hermes config path and confirm that the value is the string none, not empty or null;
  4. start a new session and run /reasoning again;
  5. if the route or provider exposes a server-side request record, echoed value, or explicit confirmation, compare the sent/mapped value with what was received or accepted; if the log contains only model, token, or spend data, record that the accepted effort cannot be confirmed instead of inferring it.

If the model requires reasoning, the off request may never be available. Use the lowest level that the route supports instead of repeatedly changing the display toggle.

Using an OpenAI-compatible custom provider

The effective behavior is shared between Hermes, the model, and the provider. As one example, the current BetterToken Hermes setup guide tells users to configure their own API key, the https://www.bettertoken.ai/v1 Base URL, and the exact model ID from the catalog, then verify the connection with a short request before wider use. BetterToken does not guarantee that every model supports every reasoning level or that a higher level improves every task.

Whichever provider you use, put the model ID, request record, and actual usage in the same comparison table. Otherwise, an unnoticed model, route, or billing-path change can be mistaken for a reasoning-effort effect.

Choose the lowest level that reliably clears your quality bar

Keep medium as the global baseline, use session settings for occasional hard tasks, and add a high, xhigh, or higher per-model override only after repeated same-task comparisons show a stable gain. For mechanical work, lower a model to low, minimal, or none only when quality stays above your threshold and measured latency or real usage improves as expected. Verify a sends mapping or provider acceptance/execution record when the route exposes one; otherwise state the evidence limit instead of treating token or billing data as proof of the accepted tier.

The useful default is not the theoretically strongest level. It is the lowest level that consistently reaches your quality target on your model and route with acceptable latency and real usage.

Ready to optimize your LLM workflow?

Join thousands of developers building faster, smarter, and more cost-effective AI applications with BetterToken.

Get Started for Free