Skip to content

Reasoning and thinking

Start with a recipe for your model; you do not need to understand every protocol field first.

  • To change thinking effort: find your connection type in the table below and adjust its supported fields.
  • If a request fails after a tool call because thinking is missing: use the replay-contract guidance below to determine whether historical thinking must be preserved.
  • To read a translation: configure thinking translation. It affects display only, not the model’s thinking request settings.

Thinking affects cost: newly generated thinking counts as output, and thinking sent again with history counts as input. See the model field reference for full field definitions.

Wire family Chord keys What comes back
Responses (type: responses) reasoning.effort, reasoning.summary Plaintext reasoning_text plus encrypted reasoning items
Chat Completions (type: chat-completions) reasoning.effort, plus family-specific fields sent through compat.request_overrides.body (thinking, enable_thinking, reasoning_split, clear_thinking, …) Usually reasoning_content; some backends inline tagged thinking in content
Messages (type: messages) thinking.type, thinking.budget_tokens, thinking.effort, thinking.display Signed thinking blocks
Gemini (type: generate-content) thinking.level, thinking.budget, thinking.include_thoughts Thought summaries plus thought signatures

reasoning.effort has no local whitelist: Chord passes the value through and lets the backend accept, clamp, or reject it. Only the Responses wire normalizes whitespace and casing first, so high and High both work there.

On Chat Completions, a gateway that translates the call into the model’s native API receives the thinking settings as that API’s own field: Gemini as extra_body.google.thinking_config, Claude as thinking: {type, budget_tokens}, DeepSeek / GLM / Kimi K2.x / Doubao as thinking: {type}, and Qwen as enable_thinking. The shape comes from compat.chat_completions.native_thinking, and only a DeepSeek route (a DeepSeek model ID or compat.reasoning_continuity.contract: deepseek) selects it on its own. The selector is also what identifies a gateway model as Gemini or Claude: without it Gemini thought signatures are not written back, and Gemini 3 rejects the request that follows each tool call (HTTP 400). See Thinking behind a Chat Completions gateway.

Chord applies the DeepSeek contract automatically when the final component of a model ID (any provider prefix removed, case-insensitive) names a DeepSeek API model:

  • deepseek;
  • deepseek-v<N> with N of 4 or later, such as deepseek-v4-pro, deepseek-v4.1-flash, or deepseek-v5;
  • deepseek-flash, deepseek-pro, deepseek-chat, or deepseek-reasoner, alone or followed by -….

A name containing distill never matches. Open-weight releases hosted by third parties under their own contracts — deepseek-r1-distill-*, deepseek-coder-*, deepseek-v3, deepseek-r1 — are not recognized either; a route that serves one of them under DeepSeek’s contract opts in with compat.reasoning_continuity.contract: deepseek.

Chat Completions and Messages preserve the selected effort across tool calls, even when an assistant returns no reasoning. max is an effort level; max_tokens limits output. DeepSeek ignores budget_tokens as a thinking-effort control.

For Messages, use thinking.type: enabled with thinking.effort; Chord sends thinking: {type: enabled} and output_config.effort. For Chat, use reasoning.effort; Chord sends reasoning_effort and thinking: {type: enabled}. thinking.type: disabled explicitly turns thinking off.

Both DeepSeek paths replay reasoning from the entire retained history, including previous user turns; generic reasoning_replay: current_turn / none settings do not shorten this window. Explicitly setting either value produces a warning at startup and in chord doctor config; remove the setting or use all to clear it. Same-target Messages blocks retain their original text and usable signatures. A replay rejection does not trigger lower effort, removal of required reasoning, or textification of tool history; the error follows the model-pool handling rules. Third-party gateways must support this contract; a third-party route whose model ID names a DeepSeek API model but serves another backend opts out with compat.reasoning_continuity.contract: none, on Chat and Messages alike.

For targets other than the DeepSeek Chat/Messages paths above, the answer depends on whether the backend requires its own reasoning content back:

  1. No thinking: the model does not reason, or you never turn thinking on. Nothing to configure.
  2. Thinking comes back, but the backend does not require it again: the default is enough. Chord replays chat-native reasoning optimistically on the first attempt and degrades to structured completed tool facts if the target rejects it. If the backend never returns reasoning_content at all, Chord reads it as replay-incompatible and drops per-request reasoning controls for the rest of the turn; set compat.chat_completions.keep_reasoning_effort: true only for endpoints that accept those controls without a reasoning-content contract (Grok on Chat Completions is the documented case).
  3. The backend validates the replayed reasoning: set compat.reasoning_continuity.mode: openai_visible plus reasoning_replay: all so every assistant message goes back unchanged. This is the contract for Kimi K3, Qwen preserve_thinking, GLM clear_thinking: false, and Xiaomi MiMo. Third-party relays do not reliably follow the official API here (some reject replayed reasoning_content, some only lose quality), so verify the actual endpoint before relying on all. The current turn’s tool loop is always required; completed turns are the optional part: keeping them improves continuity and cache reuse, dropping them saves input on every request.
  4. Responses, Messages, and Gemini: native continuity is automatic. Chord captures the plaintext or signed/encrypted state and replays it where the wire allows; nothing to configure. On the native Gemini endpoint, a model ID starting with gemini-3 also turns on the missing thought-signature repair. Behind a Chat Completions gateway, Gemini and Claude state is only replayed once compat.chat_completions.native_thinking names the family (gemini-3 for Gemini 3).

reasoning_replay: all replays completed-turn thinking on every request, which the backend bills as input. The default (current_turn) strips completed turns and is enough unless the contract in step 3 applies. Stale reasoning can also anchor the model to an outdated approach, so keep all only where the backend asks for it.

Portable visible reasoning is converted into the target’s structured carrier when one exists (openai_visible on Chat Completions, anthropic_unsigned on verified Messages-compatible endpoints); otherwise it is dropped rather than pasted into assistant content. Completed tool calls and their results stay structured, and that is the part which must survive a provider switch. See Cross-protocol fallback continuity.

  • Thinking tokens are output tokens; replayed reasoning is input tokens.
  • Completed-turn reasoning is stripped by default to keep requests small. Anthropic filters prior-turn thinking blocks server-side and bills only the blocks the model actually sees, so omitting them costs nothing there; backends that replay history verbatim bill every retained token, which is why the contracts above opt into all explicitly.
  • Some backends pin sampling or invalidate caches when thinking settings change: Kimi K3 fixes temperature / top_p / penalties and drops the prefix cache when reasoning_effort changes mid-session.
  • Thinking translation for the TUI (thinking_translation) is display-only and is never written back into model context; see Appended thinking translation.

When a request fails with a thinking-mode error, start from Troubleshooting.