Skip to content

Configuration & Auth

Connect your models once, then reuse pools, fallback, and project overrides. Chord separates behavior configuration and credentials:

  • ~/.config/chord/config.yaml: providers, models, extensions, defaults
  • ~/.config/chord/auth.yaml: API keys / OAuth credentials
  • .chord/config.yaml: project-level overrides
  • ~/.config/chord/agents/ and .chord/agents/: agent role definitions

You do not need to read this page from top to bottom:

A practical precedence model is:

  1. Built-in defaults
  2. Global config
  3. Project config
  4. Agent-level config

This lets you keep personal defaults, project-specific behavior, and per-agent capabilities separate.

Project configuration is read from .chord/config.yaml in the startup directory; Chord does not search parent directories. Set only the fields you want to override and leave the rest inherited. Merge rules:

  • omitted project fields stay truly unset instead of silently shadowing global defaults;
  • unrecognized keys, wrongly typed values, and out-of-range settings in any config file are logged to chord.log and treated as not configured while the rest of the file still applies; in a project config an invalid leaf falls back to the inherited global value. Malformed YAML (a syntax error) prevents startup; run chord doctor config for the full problem list;
  • settings that load as written but likely do not work as intended, such as a thinking block a Chat Completions gateway never receives, are logged to chord.log as warnings and listed by chord doctor config without counting as problems;
  • global-only keys such as paths.* and maintenance.* (also model_templates and diagnostics) are ignored in project config;
  • most scalar and object values override the global value at the same key;
  • model_pools merge by pool name, with same-name project pools overriding the global definition;
  • mcp merges by server name, with each same-name project server replacing the entire global server definition rather than inheriting individual connection or permission fields;
  • append-style extension points keep global entries and add project entries: currently skills.paths and per-trigger hook arrays under hooks.* append rather than replace.

If global config.yaml is missing, the first chord run starts a one-time setup wizard that writes config.yaml and, when needed, auth.yaml; see Quickstart. Prefer to write both files yourself? Everything below is the full field reference.

providers:
openrouter:
type: chat-completions
api_url: https://openrouter.ai/api/v1/chat/completions
models:
openai/gpt-5.5:
limit:
context: 400000
input: 272000
output: 128000
modalities:
input: [text, image, pdf]

Chord’s api_url is the complete request URL. For the BigModel Coding Plan OpenAI-compatible endpoint, append /chat/completions to the Coding Plan base.

providers:
bigmodel:
type: chat-completions
api_url: https://open.bigmodel.cn/api/coding/paas/v4/chat/completions
models:
glm-5.3: &bigmodel-glm-5-3
limit:
context: 1000000
output: 128000
reasoning:
effort: max
compat:
request_overrides:
rename_body_fields:
max_completion_tokens: max_tokens
body:
thinking:
type: enabled
clear_thinking: false
reasoning_continuity:
mode: openai_visible
glm-5.3-flash:
<<: *bigmodel-glm-5-3
modalities:
input: [text, image, pdf]

glm-5.3 is the text-only flagship; glm-5.3-flash takes the same text parameters and settings and adds image/PDF input, so pick the model ID your workload needs.

For provider/model-specific copy-paste snippets (GPT-5.4/5.5/5.6, Claude, Gemini, GLM, DeepSeek/OpenAI-compatible), see Model configuration recipes.

providers:
openai:
type: responses
api_url: https://api.openai.com/v1/responses
models:
gpt-5.6-sol:
limit:
context: 1050000
output: 128000
reasoning:
effort: medium
summary: auto
variants:
low:
reasoning:
effort: low
high:
reasoning:
effort: high
xhigh:
reasoning:
effort: xhigh
max:
reasoning:
effort: max
modalities:
input: [text, image, pdf]
model_pools:
default:
- openai/gpt-5.6-sol@xhigh

Pair this provider with an API key in ~/.config/chord/auth.yaml:

openai:
- "$OPENAI_API_KEY"
  • gpt-5.6-terra and gpt-5.6-luna are wired the same way; keep the models key and the model_pools ref on the same model ID.
  • This snippet targets the official OpenAI API, so it declares the full 1050000 window with no input: Chord then derives the usable input budget as context minus the model’s own output cap (1050000 − 128000 = 922000), and reserves the default 64000 output cap only for models that declare no limit.output. Above 272K is a pricing threshold here, not an input cap, so do not add input: 272000.
  • Codex OAuth uses the same model windows as the API: GPT-5.4 / 5.6 / 6 run the 1050000 / 922000 / 128000 allocation there too (see OpenAI Codex preset below).
  • Supported API reasoning efforts are none, low, medium, high, xhigh, and max; select a configured variant with a ref such as openai/gpt-5.6-sol@xhigh.
  • When reasoning is active, Responses defaults reasoning.summary to auto; set it to none to opt out explicitly. Chord does not currently expose GPT-5.6 reasoning.mode: pro.
  • preset: codex providers can also use max when the selected model/backend supports it. Whether a given effort level is accepted is model/provider-specific.

Limits and reasoning values were checked against the current Codex model catalog and OpenAI’s GPT-5.6 model guidance. Actual relay allocations may be smaller and should follow the gateway’s published model catalog.

Read model limits in this order:

  1. limit.context is the total window. For most models, input + requested output just needs to fit inside this number.
  2. limit.input is only needed when the provider also lists a separate input cap. Some GPT models work this way; if you omit it, Chord derives the usable input budget as limit.context minus the model’s own limit.output (only a model declaring no output cap falls back to the global max_output_tokens default). A declared limit.input is always used as-is.
  3. limit.output is the model’s own output capacity. Chord’s default requested output cap (max_output_tokens) is 64000, so real requests use min(64000, limit.output) before the available-context clamp. Set max_output_tokens explicitly to choose a different global cap. If a model’s real output capacity is below 64000 and limit.output is omitted, backends that validate the requested max_tokens server-side will reject those requests. Declare limit.output for such models, or lower the global max_output_tokens.

parallel_tool_calls defaults to true for Responses and Chat Completions providers. Set it to false on a provider, model, or variant only when the backend or workflow requires serial tool calls. Provider-level user_agent is also available for gateways that require a specific client identifier.

Provider auth headers are inferred separately from type, but can be overridden with auth_scheme when a compatible endpoint expects a different credential header:

  • type: messages → default auth_scheme: anthropic-api-key (x-api-key)
  • type: responses → default auth_scheme: bearer (Authorization: Bearer)
  • type: chat-completions → default auth_scheme: bearer

For an Azure OpenAI Responses endpoint, configure a plain type: responses provider with auth_scheme: api-key and store: true, and remove the Codex identity headers with compat.request_overrides.headers set to null (see below).

Supported auth_scheme values are:

  • anthropic-api-key
  • bearer
  • api-key

Use an explicit override only when the endpoint’s auth requirements differ from Chord’s transport default. For example, a provider may expose an Anthropic-compatible /messages path but require Authorization: Bearer instead of x-api-key. In that case, keep the same type and set only auth_scheme: bearer.

For Anthropic’s gated 1M context beta, Chord opts in only when the model declares a window of at least 1M tokens (limit.input when set, otherwise limit.context). Models with smaller declared windows do not receive the beta header. The provider may apply different access requirements and pricing above 200K tokens.

store controls whether a Responses backend retains requests and responses server-side. It defaults to false. Enable it only when the backend explicitly requires server-side retention and you accept the data-retention trade-off. Do not enable it for preset: codex; the official Codex OAuth endpoint rejects store: true.

Codex OAuth uses the same model windows as the API examples.

providers:
codex:
preset: codex
type: responses
models:
gpt-5.5:
limit:
context: 400000
input: 272000
output: 128000
gpt-5.4:
limit:
context: 1050000
input: 922000
output: 128000
gpt-5.6-sol:
limit:
context: 1050000
input: 922000
output: 128000

GPT-5.4 / 5.6 Sol / Terra / Luna / GPT-6 Sol / Luna / Astra / GPT-6.1 Sol use 1050000 / 922000 / 128000 (1.05M total window; the 922K input budget derives as context − output, since these models publish no separate input cap); GPT-5.5 and GPT-5.2 use 400000 / 272000 / 128000. See Model configuration recipes for complete examples.

preset: codex can use OpenAI / ChatGPT OAuth credentials from auth.yaml. OAuth entries are mappings:

codex:
- refresh: rfr_...
access: eyJ...
expires: 1774009702606
account_id: acc_... # optional; only workspace/account tokens always have this
account_user_id: u_...__acc_... # optional; parsed/backfilled in the background when missing
email: user@example.com # optional

account_id is not present for every ChatGPT account. Personal Plus/Pro access tokens may carry only user_id and no chatgpt_account_id; Chord still uses those credentials as ordinary OAuth bearer tokens and omits the ChatGPT-Account-ID header. Features that require a workspace/account id, such as Codex usage / rate-limit polling, skip those credentials until an account id is provided or parsed later.

For large account pools, Chord does not synchronously parse every OAuth JWT at startup and does not block provider initialization when one access token lacks account_id. Startup reads only metadata already present in auth.yaml; missing account_user_id, account_id, email, and expires are parsed and backfilled in the background after the provider is available. When manually converting Codex / sub2api / other login exports, keep any available account_id, account_user_id, and email, but they are not startup requirements.

Configure an Azure OpenAI Responses endpoint as a plain type: responses provider with auth_scheme: api-key and store: true. Omit the Codex streaming identity headers with compat.request_overrides.headers set to null:

providers:
azure:
type: responses
api_url: https://YOUR-RESOURCE.openai.azure.com/openai/v1/responses
auth_scheme: api-key
store: true
trust_http_400: true
retry_after_max_s: 86400
compat:
request_overrides:
headers:
OpenAI-Beta: null
originator: null
models:
gpt-5.5:
limit:
context: 400000
input: 272000
output: 128000

Azure’s v1 Responses endpoint can use /openai/v1/responses directly; add api-version only when you need to pin or opt into a specific version such as preview. Store the Azure API key under the same provider name in auth.yaml:

azure:
- $AZURE_OPENAI_API_KEY
providers:
gemini:
api_url: https://generativelanguage.googleapis.com/v1beta/models
models:
gemini-3.8-flash:
limit:
context: 1048576
output: 65536
modalities:
input: [text, image, pdf]

For Gemini, set api_url to the /models base path. Chord detects type: generate-content from the URL path’s /models suffix, so type can be omitted. Do not include the model name or :streamGenerateContent?alt=sse; Chord appends /{model}:streamGenerateContent?alt=sse automatically. The model map key, such as gemini-3.8-flash, is the model ID sent to Gemini.

Gemini thinking options use the same unified thinking object as other providers (no separate gemini_thinking key):

  • thinking.budget → generationConfig.thinkingConfig.thinkingBudget
    • Gemini: ✅ used
    • Anthropic: ⚠️ only when thinking.type: enabled (mapped to Anthropic budget mode)
    • OpenAI: ❌ ignored
  • thinking.include_thoughts → generationConfig.thinkingConfig.includeThoughts
    • Gemini: ✅ used
    • Anthropic / OpenAI: ❌ ignored
  • thinking.level → generationConfig.thinkingConfig.thinkingLevel (minimal|low|medium|high, Gemini 3+; not all models support minimal)
    • Gemini (3+): ✅ used
    • Gemini 2.x / Anthropic / OpenAI: ❌ ignored

Example:

providers:
gemini:
api_url: https://generativelanguage.googleapis.com/v1beta/models
models:
gemini-2.5-flash:
limit:
context: 1048576
output: 65536
modalities:
input: [text, image, pdf]
thinking:
budget: -1
include_thoughts: true
gemini-3-pro:
limit:
context: 1048576
output: 65536
modalities:
input: [text, image, pdf]
thinking:
budget: -1
level: high

If type is omitted, Chord auto-detects it from provider config:

  • preset: codex → responses
  • api_url path ending in /responses → responses
  • api_url path ending in /chat/completions → chat-completions
  • api_url path ending in /messages → messages
  • api_url path ending in /models → generate-content

If none of these rules match, set type explicitly.

If your model outputs English thinking / reasoning and you want an appended translation (for example, Chinese) in the TUI, you can enable thinking_translation:

model_pools:
translation:
- openai/gpt-5.4-mini
thinking_translation:
target_language: zh-Hans
model_pool: translation
max_chars: 1000

Notes:

  • Only thinking / reasoning output is translated, never the assistant final answer. The translation is appended under the corresponding thinking card with a neutral Translated · <target_language> header, rendered through the same Markdown / code-highlighting pipeline, and never written back into model context.
  • target_language and model_pool are both required; if either is missing, the feature is disabled. model_pool must point to a top-level model_pools entry: prefer a separate low-cost translation pool. The pool can contain multiple provider/model[@variant] refs; translation runs a single fallback round across them in order, moving to the next candidate on failure (including network/5xx/timeout) or when a result is empty, clearly truncated, or in the wrong language.
  • max_chars (default 1000) limits the thinking preview sent for translation; only the leading max_chars runes are translated, and text past that prefix will not appear in the translated card. Set a smaller value such as 500 for lower latency/cost, or a larger one for more complete translations.
  • A temporary failure only skips that one thinking block; it does not block later thinking translations or the main response. Per-provider transport timeouts (one-minute-class by default) still apply, so a stalled model or key can fail over while the rest of the pool gets a chance to run.
  • Translations are persisted in the session directory (thinking_translations.json) and restored when the session is resumed. A given thinking block is translated at most once: changing thinking_translation.target_language later does not re-translate already-stored blocks.

More detailed fields are described in the config reference below.

Provider keys must match the provider name in config.yaml:

The first-run wizard can create this file for you. It supports either literal API keys or $ENV_VAR placeholders.

anthropic:
- "$ANTHROPIC_API_KEY"
openai:
- "$OPENAI_API_KEY"

You can list multiple keys for rotation or backup.

For preset: codex OAuth providers, Chord keeps frequently changing runtime status (quota snapshots, reset times, last warm-up timestamps, shared OAuth status cache) in auth.state.json, not in auth.yaml.

That split is intentional:

  • auth.yaml remains the user-edited source of truth for credentials and stable OAuth fields such as refresh, access, expires, account_id, and email; empty OAuth fields are omitted when Chord rewrites the file, and OAuth status does not belong in auth.yaml;
  • auth.state.json is machine-managed shared runtime state. Normal entries are keyed directly by account_user_id below each provider so quota / reset updates and account states such as expired, deactivated, and invalidated do not constantly rewrite auth.yaml while the user may also be editing it. Refresh-only credentials whose account is not known yet can temporarily use a refresh_sha256:<digest> state entry until the first successful refresh backfills account_user_id. State entries without a matching auth.yaml OAuth credential, and unrecognized legacy state-key formats, are removed by chord auth state clean.

For OAuth credentials with access, the access token must carry parseable account and user/account-user claims. If auth.yaml already has account_id, the token’s account ID must match; otherwise the access token is rejected as a mismatched credential. Chord can also keep a refresh-only OAuth entry (refresh without access) and refresh it on first use; after a successful refresh, Chord extracts account_id and switches runtime state to the account_user_id key.

If refresh fails unrecoverably before the account is known, Chord records the invalid state under refresh_sha256:<digest> so chord auth state clean can remove the unusable credential later. An OAuth entry with neither access nor refresh is unusable.

expires is the access-token expiry timestamp in Unix milliseconds. When access contains a JWT exp claim, Chord uses that value as the most accurate expiry metadata and can cache the resulting expiry in auth.state.json without storing the access token there. A missing or locally expired expires value does not by itself mark an OAuth slot expired or unhealthy. Chord still tries the existing access token first, and only after an authentication failure will it refresh the credential or mark it expired if recovery is impossible.

Typical auth.state.json content looks like:

{
"openai": {
"user-1__acc-1": {
"account_user_id": "user-1__acc-1",
"account_id": "acc-1",
"email": "user@example.com",
"expires": 1774009702606,
"status": "expired",
"updated_at": 1774009702606,
"last_warmup_at": 1774009702606,
"codex_primary_used_pct": 12.5,
"codex_primary_window_minutes": 60,
"codex_primary_reset_at": 1774013302000,
"codex_secondary_used_pct": 40,
"codex_secondary_window_minutes": 10080,
"codex_secondary_reset_at": 1774600000000
}
}
}

The status field is authoritative only in auth.state.json. Chord writes expired when an access token can no longer be used and the credential cannot be refreshed (including missing, invalid, expired, or reused refresh tokens), deactivated when the service reports a disabled/banned account, and invalidated when the account must be re-authenticated. Any non-empty status makes that OAuth slot unselectable until it is cleaned up or replaced.

These cached Codex quota/reset fields are restart-stable scheduling and display hints, not hard blocks by themselves:

  • they help startup / first-pick ordering choose accounts that are more likely to still have quota;
  • they let key switches immediately show the last cached snapshot before a fresh warm-up completes;
  • they do not by themselves make the account absolutely unselectable;
  • real hard blocking still comes from confirmed request failures and runtime cooldown state.

Provider credentials in auth.yaml support environment-variable expansion for scalar API-key values:

anthropic:
- "$ANTHROPIC_API_KEY"
openai:
- "${OPENAI_API_KEY}"

Expansion is applied when the scalar starts with $. Unset variables expand to an empty string and are filtered out, unless the YAML value is a literal empty string. This expansion applies to auth.yaml credentials, not generally to every field in config.yaml.

If you intentionally need an empty API key, write a literal empty string:

local-provider:
- ""

Do not rely on an unset environment variable for this case. An unset $ENV_VAR is treated as a missing credential and is filtered out.

When a provider has multiple API keys / OAuth accounts, Chord uses two settings: key_rotation controls when Chord reselects a key, and key_order controls how Chord chooses among selectable keys.

  • key_rotation: on_failure (default): keep using the current key until it fails, cools down, or becomes unusable.
  • key_rotation: per_request: reselect a key before every request; useful for load balancing across independent keys.
  • key_order: sequential (default for non-Codex providers): choose in stable key order, generally preferring the least-recently-used selectable key.
  • key_order: random: choose randomly among selectable keys.
  • key_order: smart: Codex providers only. Prefer healthy OAuth accounts with better quota headroom and reset timing.

key_rotation only rotates credentials / API keys. It does not rotate models; model selection still follows the model pool sticky cursor and fallback logic.

Loop mode still follows the configured key_rotation / key_order. For Codex long-running loops, keep the default key_rotation: on_failure if you prefer stable transport/cache continuity; explicitly use per_request only when you want to distribute quota across multiple accounts.

Only providers with preset: codex are treated as OAuth providers.

For Codex providers, prefer configuring only preset: codex plus model settings. Do not manually override preset-managed fields such as api_url, token_url, client_id, type, store, responses_websocket, or supported_service_tiers unless you are deliberately testing transport internals. The preset selects the official OAuth transport, Responses endpoint, WebSocket/cache defaults, quota polling, smart key ordering, and service-tier capability.

It does not define a separate HTTP request body or force a Codex User-Agent: non-Codex type: responses providers use the same Responses wire shape described above, and all providers default to User-Agent: chord/<version>. Use supported_service_tiers when you need an explicit tier matrix.

Codex OAuth account selection is controlled by key_rotation / key_order in Provider key selection. Codex defaults to key_order: smart, which considers quota snapshots, soft cooldown, and reset timing when choosing an account.

smart ranks selectable Codex OAuth accounts by preferring:

  • accounts whose cached snapshot has no tracked window at 100% used; a 100% window is tried last but is not a hard block by itself;
  • accounts with remaining quota in the shorter primary window (for example the 5h window), choosing the nearer primary reset first so soon-expiring quota is used before it is wasted;
  • then accounts with remaining quota in the longer secondary window (for example the 1w window), again preferring the nearer secondary reset;
  • then higher remaining headroom when the comparable windows reset at the same time;
  • still falling back to unknown / stale candidates when no better option exists.

When a Codex client becomes active, Chord may also background-probe additional OAuth slots to refresh cached headroom snapshots. That warm-up is best-effort, low-concurrency, cancels when the active client is replaced, and only refreshes cached quota state; authentication failures from usage probes do not mark OAuth credentials unusable.

Warm-up priority is also state-aware:

  • OAuth slots that have never been warmed up in shared state are probed first;
  • older cached entries are refreshed before recently refreshed ones;
  • after warm-up or polling returns a newer snapshot, Chord writes it to auth.state.json and other processes adopt it lazily when they next read key-selection or rate-limit state.
Terminal window
# auto-select a configured codex provider
chord auth
# explicitly choose a provider
chord auth codex
# headless / SSH environments
chord auth codex --device-code

Chord selects the active model via named model pools. Each pool entry should be a full provider/model[@variant] reference so the provider endpoint, auth, protocol, and variant tuning are unambiguous.

Pool definitions live in config.yaml (global or project-level). Agent configs may reference pool names to restrict access; they cannot define inline pools.

# ~/.config/chord/config.yaml or .chord/config.yaml
model_pools:
thinking:
- anthropic/claude-opus-5
- openai/gpt-5.5
non-thinking:
- anthropic/claude-sonnet-4

Project-level .chord/config.yaml model_pools are merged into the global config (same-name pools override).

Agents do not need to set model_pools. If omitted, the agent can use every pool defined in merged config.yaml model_pools, sorted by pool name. Add model_pools: [...] only when you want to restrict that agent to a subset or customize its fallback order.

# ~/.config/chord/agents/builder.yaml or .chord/agents/builder.yaml
name: builder
mode: main
model_pools: [thinking, non-thinking]
.chord/agents/reviewer.yaml
name: reviewer
mode: subagent
model_pools: [thinking]

When no pool is explicitly selected, Chord falls back to the agent’s first allowed pool: the first entry in model_pools: [...] when configured, otherwise the alphabetically first top-level pool.

At runtime, use /models to switch the pool for the current view (per project, persisted across restarts). In the main view this means the current main role; in a SubAgent view it means that SubAgent’s agent pool selection.

Switching pools updates the full fallback chain for subsequent LLM calls, even if the currently selected provider/model exists in both pools (in-flight requests keep using their starting snapshot). You can also set a named agent directly with /models --agent <name> <pool>. For SubAgents, the default behavior is to use the first allowed pool; switching back to that pool restores the default behavior.

User messages submitted while the main agent is busy remain queued until a safe request boundary. If the current provider/model attempt fails and Chord is about to send a request to the next fallback model, messages queued by that point are committed to the conversation and included in that fallback request. Chord never rewrites a provider request that is already in flight; messages arriving after the fallback request starts wait for the next request boundary.

Reusing protocol templates with YAML anchors

Section titled “Reusing protocol templates with YAML anchors”

Chord has no model_templates schema field. You can still use YAML anchors and merge keys under that top-level container; Chord ignores the container itself and reads the expanded model entries under providers.

Merge keys (<<:) copy the referenced mapping into the current entry at the key level, and the current entry wins on conflict:

  • Scalar fields (reasoning.effort, a compaction fraction) simply replace the inherited value.
  • Nested objects are replaced as a whole, not merged field by field: overriding with limit: {context: 1050000} on a template that already declared limit: {context: 1000000, output: 128000} silently drops the inherited output. Write the complete block for anything you override (limit, cost, compaction, variants entries, modalities, …).
  • The same rule holds along a chain (gpt-5.6-luna: &gpt-5-6-luna {<<: *gpt-5-6-base}): the deepest entry wins per whole key.
  • You cannot unset a field an ancestor template declares: override it with a concrete value, or stop referencing that template. compaction.threshold: 0 and compaction.reminder: -1 are the documented exceptions that disable those two behaviors explicitly.

This page covers protocol and field semantics. For current model limits, pricing, and complete GPT / Claude / Gemini / GLM / DeepSeek snippets, see Model configuration recipes.

model_templates:
chat-thinking: &chat-thinking
limit:
context: 200000
output: 64000
reasoning:
effort: high
compat:
# DeepSeek and GLM Chat APIs document max_tokens, not the OpenAI
# reasoning-model default max_completion_tokens.
request_overrides:
rename_body_fields:
max_completion_tokens: max_tokens
body:
# Gateways for the model families listed in "Thinking behind a Chat
# Completions gateway" (model-configs.md) get this object written
# automatically; the override covers other backends — and any extra
# field of the same object, such as GLM's clear_thinking.
thinking:
type: enabled
# Replays native reasoning_content and accepts portable visible
# reasoning from other wire families without injecting request fields.
reasoning_continuity:
mode: openai_visible
responses-thinking: &responses-thinking
limit:
context: 200000
output: 64000
reasoning:
effort: high
summary: auto
messages-thinking: &messages-thinking
limit:
context: 200000
output: 64000
thinking:
type: adaptive
effort: high
compat:
# Compatible endpoints should not receive Anthropic beta headers unless
# their own documentation opts into them.
request_overrides:
headers:
anthropic-beta: null
providers:
chat:
type: chat-completions
api_url: https://example.com/v1/chat/completions
models:
chat-model: *chat-thinking
responses:
type: responses
api_url: https://example.com/v1/responses
models:
responses-model: *responses-thinking
messages:
type: messages
api_url: https://example.com/v1/messages
models:
messages-model: *messages-thinking

Model field semantics:

  • limit.context: total request window when the provider publishes one. If it is omitted and both limit.input and limit.output are positive, Chord derives it as input + output; an explicit context always takes priority.
  • limit.input: independent input cap when published. A declared value is authoritative and used as-is, even when it is not additive with limit.output inside the window. If omitted, Chord derives the prompt budget as limit.context minus the model’s limit.output (so the 1.05M GPT family gets 1050000 − 128000 = 922000); only a model declaring no output cap falls back to reserving the effective default output cap (max_output_tokens, default 64000).
  • limit.output: model output capacity. Runtime requests are also capped by the global max_output_tokens setting and remaining total-context space.
  • reasoning.effort: reasoning depth. Chord keeps no local whitelist: whatever level the provider supports reaches the upstream unchanged, and the Responses wire additionally normalizes whitespace and casing before sending.
    • Chat Completions sends top-level reasoning_effort.
    • Responses sends reasoning.effort and optional reasoning.summary.
  • reasoning.effort_map: maps the canonical effort value to the wire value the provider actually accepts, for example {high: max} when a gateway exposes max for Chord’s high. The mapping applies to the final resolved effort, so variant-level maps replace the model-level map for that variant.
  • reasoning.summary: Responses reasoning summary request. Supported Chord values are auto, concise, detailed, and none. When reasoning is active, omission defaults to auto so cross-provider replay retains portable summary text; use none to opt out explicitly.
  • thinking: Messages-compatible extended thinking. type: adaptive combines with thinking.effort, which Chord sends as output_config.effort.
  • text.verbosity: optional OpenAI-compatible visible-text verbosity hint.
  • variants: named model parameter overrides selected with refs such as provider/model@high.
  • cost: optional USD-per-million-token estimates. It can include input/output, cache prices, service-tier multipliers, and long-context input tiers.
  • modalities.input: supported input kinds: text, image, and pdf.
  • supported_service_tiers: accepted non-standard tiers such as fast or slow; price multipliers are configured separately under cost.

Compatibility fields:

  • compat.request_overrides.body: recursively merges arbitrary JSON into the final protocol request. A null value deletes that field.

  • compat.request_overrides.rename_body_fields: renames a final request field while preserving Chord’s dynamically computed value. Use this for differences such as max_completion_tokens: max_tokens.

  • compat.request_overrides.headers: sets arbitrary request headers. A null value removes a Chord default header, for example anthropic-beta: null on a compatible Messages endpoint.

  • compat.reasoning_continuity.mode:

    • none: no provider-specific visible reasoning replay.

    • openai_visible: replays unchanged assistant reasoning_content during Chat Completions tool loops and accepts portable visible reasoning from other wire families as reasoning_content. It does not inject request fields; configure those with request_overrides.body. On the first attempt Chord still optimistically replays chat-native reasoning to any Chat Completions target, even across providers, so documented in-provider upgrades such as Kimi K2.6/K2.7 to K3 and same-model provider fallback both keep continuity.

      If a target rejects that request, Chord degrades the target for the rest of the session while keeping completed tool calls and paired results structured until strict compatibility requires text facts. On openai_visible Responses targets whose thinking mode requires replayed function-call turns to carry reasoning, Chord replays the native plaintext reasoning_text when available.

      For turns that lost their native reasoning (for example after a cross-provider model switch), a backend rejection escalates the replay to plain-text historical tool records, so the continuation no longer needs the missing reasoning. For third-party OpenAI-compatible gateways, Chord keeps the configured endpoint and does not redirect to DeepSeek’s official /beta endpoint.

      A reasoning-only output truncation may receive one bounded request-only reasoning replay on that same endpoint; if the gateway rejects it, Chord falls back to the ordinary recovery prompt.

    • anthropic_unsigned: opt-in for Messages-compatible models such as DeepSeek/GLM endpoints that return visible thinking blocks without Claude signatures. Unsigned thinking is replayed natively only to the same provider/model on the first attempt. For compatible targets, portable visible reasoning from other wire families is converted into unsigned thinking blocks instead of being injected into assistant text.

    • Responses, signed Claude Messages, and Gemini otherwise use their protocol-native continuity mechanisms automatically. Chord captures opaque encrypted/signature state and replays it only where the target wire and provenance permit. Across incompatible protocols, portable visible reasoning is converted only when the target has a structured carrier (openai_visible or anthropic_unsigned); otherwise it is dropped. Opaque state is never fabricated, and reasoning is never injected into ordinary assistant content.

      Completed tool facts are converted to the target protocol’s structured representation whenever possible and textified only if a target rejects that shape. The achieved degradation level is remembered per target.

  • compat.reasoning_continuity.reasoning_replay: selects how much reasoning Chord replays from completed turns (everything before the last user message). DeepSeek Chat/Messages always preserve the complete reasoning history; for other targets the default current_turn strips completed-turn reasoning, all replays it unchanged, and none also strips the current turn.

    Completed turns are stripped by default to keep the request small. Anthropic filters prior-turn thinking blocks server-side and bills only the blocks the model actually sees, so omitting them costs nothing there; targets that replay history verbatim charge for every retained token, which is why the contracts listed below opt into all. The policy covers every reasoning payload — plaintext reasoning_content, unsigned thinking blocks, and provider-bound native thinking (signed Claude blocks, Responses reasoning items, Gemini thought signatures) — while the tool trajectory, including call/result pairing, always survives the strip.

    Set all when the target’s contract requires the complete assistant history (Kimi K3 and keep: all models, Qwen preserve_thinking, GLM clear_thinking: false, Xiaomi MiMo; DeepSeek Chat/Messages already keep it without this setting); historical reasoning is then replayed unchanged and billed on every request. Set none only for an endpoint you verified accepts a request without the current turn’s reasoning. Current-turn reasoning — everything after the last user message, including its tool loop — is otherwise replayed unchanged.

    Anthropic additionally binds each thinking block to the conversation prefix that produced it: when a history rewrite invalidates that binding and the API rejects the replay with an invalid-signature error, Chord retries once with the thinking blocks dropped and keeps the turn’s text and completed tool facts.

    Request-scoped turn overlays (per-turn <system-reminder> hints) are not counted as user turns, so an overlay appended at the tail cannot shift the completed-turn boundary past the current turn and strip the reasoning the backend consumes in this turn’s tool chain.

  • compat.forced_tool_choice.suppress_in_thinking: downgrades loop-forced tool_choice: required to the backend default while reasoning/thinking is active. Enable it only for OpenAI-compatible endpoints that reject forced tool choice in thinking mode; ordinary tool availability and automatic tool choice still work.

  • compat.forced_tool_choice.auto_only: downgrades any non-auto tool_choice to the backend default unconditionally. Enable it for backends that only support tool_choice: "auto" and reject required, none, or named choices with a 400; automatic tool choice still works. Chat Completions, Messages, and Gemini omit the field; Responses still sends an explicit tool_choice: "auto" unless compat.responses.send_tool_choice is false.

  • compat.thinking_toolcall: enables a provider-specific parser for gateways that encode tool calls inside visible reasoning text. Leave disabled unless the gateway requires that format.

Provider-level compat values are defaults. A model-level compat block can override them for one model.

Request overrides apply to HTTP transports after Chord has constructed the protocol request. Configuring them on a Codex Responses provider disables its WebSocket transport for that request so the final JSON patch can be honored.

If a project needs local defaults, create this file at the project root:

.chord/config.yaml

Common uses include:

  • Project-specific permission rules
  • Project-specific LSP / MCP / Hooks / Skills settings

Provider-level compress selects the encoding for compressed upstream request bodies: gzip or zstd (zstd is the codec the Codex client uses for codex-backend request bodies). It is different from context management (compaction / reduction): it only changes HTTP request transfer encoding and does not summarize or remove conversation history.

providers:
openai:
compress: gzip
codex:
preset: codex
compress: zstd # codex-backend accepts zstd request bodies

Chord compresses the request body only if compression reduces the payload size; otherwise it sends the request uncompressed. Compression failures are logged and the request is sent uncompressed too. The response direction is unaffected: Chord still advertises only gzip responses and decodes them itself.

Provider/model requests identify the client with User-Agent: chord/<version> by default. Set provider-level user_agent only when a provider or gateway requires a specific value:

providers:
gateway:
user_agent: RequiredGatewayClient/1.0

This setting also applies to Responses HTTP requests, Codex OAuth requests, Codex usage polling, and Responses WebSocket handshakes for that provider. WebFetch uses its own web_fetch.user_agent.

To send a Codex-style User-Agent, set the captured Codex value explicitly and keep it updated when the Codex client version or terminal changes:

providers:
codex:
preset: codex
user_agent: "codex-tui/0.139.0 (Mac OS 15.3.2; arm64) ghostty/1.3.1 (codex-tui; 0.139.0)"

Provider-level retry settings control the generated delay between complete retry rounds. A round starts at the current sticky model-pool cursor and walks the remaining pool entries and their eligible keys without sleeping between fallback targets. The cursor provider’s settings own the delay before the next round; if a fallback later succeeds and becomes the sticky cursor, its settings apply to subsequent requests. The same settings also control ordinary HTTP 429 key cooldown when the response carries no Retry-After hint; a valid hint (bounded by retry_after_max_s) always outranks them.

providers:
gateway:
retry_backoff: exponential
retry_delay_ms: 500
  • retry_backoff: exponential is the default. It starts at retry_delay_ms (default 1000) and doubles each round, capped at 60 seconds.
  • retry_backoff: fixed uses retry_delay_ms for every round.
  • retry_backoff: none disables generated round backoff. An ordinary 429 without a Retry-After hint marks the failed key as recovering without a timed cooldown, so healthy keys and fallback targets still take priority. Rounds are unlimited by default, so with none a fast-failing upstream is retried back-to-back without delay.
  • retry_delay_ms accepts 0 through 60000; 0 / omitted means 1000ms. Invalid modes, negative values, and values above the cap are configuration errors.

A retry round that pauses before its next attempt shows that wait in the status bar as a countdown to the attempt (↺ round 12 · retry in 45s). A round with no delay re-probes immediately and keeps showing how long the retry has been running, since it has no wait to count down.

For an ordinary 429, the key cooldown follows a single priority order: a confirmed quota reset window wins, then a valid Retry-After (bounded by retry_after_max_s) applies verbatim, and only a hint-less 429 falls to the retry pacing above: the configured exponential/fixed/none mode, or the one-second exponential default when neither field is set. Invalid or deactivated credentials, and cooldowns already established by other hard states, are never shortened or cleared.

This 429 pacing applies before and after visible streaming output alike: a 429 that interrupts a visible stream cools the key down and rotates to the next one.

When every key of every pool entry is cooling down, Chord waits instead of sending a request, and how long it sleeps depends on the pool. With a single model configured there is nothing else to try, so it waits for the earliest key recovery instant — a confirmed quota reset instant the provider defines, or a Retry-After hint already capped by retry_after_max_s. Nothing re-probes the pool during that stretch, so a credential added mid-wait is picked up once the wait ends. With fallback models configured it re-checks the pool at least once a minute, because a sibling model, a newly added credential, or a refreshed rate-limit snapshot can free up a request long before the longest cooldown ends. Either way the status bar counts down to the point a request can actually go out, not to the next internal re-check. The shortest cooldown in the pool always decides: a model that is ready again is never held back by a longer cooldown on another one, and a pool that is re-checked every minute picks a recovered key up as soon as it is ready.

When a key goes into cooldown, the API failure behind it is recorded in the error panel (Ctrl+E) by the wait that shows the cooling or by the next foreground request (an agent turn or a context compaction), whichever comes first. This includes failures from background work such as memory extraction or thinking translation. A credential the provider permanently invalidated (an expired refresh token, an invalidated or deactivated account) is announced by the next foreground request. Each failure is recorded once: a failure already reported as a retry error by the attempt that caused it is not repeated.

Codex OAuth follows the same rules: every Codex 429 is an ordinary 429. A retry hint (Retry-After or WebSocket resets_in_seconds) is honored ahead of explicit settings, and a usage-limit 429 carrying neither a hint nor an exhausted quota snapshot uses the ordinary defaults above, not the one-minute cooldown of the codex preset, which still covers non-429 usage-limit errors. When a Codex rate-limit snapshot shows an exhausted window with a future reset, Chord treats that as confirmed quota exhaustion and keeps the provider reset authoritative.

Project-level .chord/config.yaml can override these fields for one provider.

Provider-level timeout settings are optional and use seconds. Unset or 0 keeps the built-in defaults.

providers:
codex:
response_header_timeout: 180
stream_idle_timeout: 90
stream_total_timeout: 1800
websocket_handshake_timeout: 45
  • response_header_timeout: timeout from starting a streaming HTTP request until response headers arrive, including connection setup and request-body upload. It stops once headers arrive and does not cap the total duration of a healthy stream; use stream_idle_timeout to bound gaps between streamed chunks. 0 keeps the built-in default.
  • stream_idle_timeout: maximum idle time between streamed model data. When set, it overrides both the normal SSE idle timeout and slow-phase idle timeout for that provider, and it also applies to Codex Responses WebSocket reads.
  • stream_total_timeout: wall-clock cap in seconds for one stream, counted from when the response body starts. When set, it also applies to Codex Responses WebSocket reads. 0 / omitted keeps the default of no cap: a stream that keeps producing data is slow, not broken, and is bounded by stream_idle_timeout alone. Set it to bound the one shape the idle timeout cannot catch: a stream that drips data often enough to reset the idle timer but never finishes. The read fails with a timeout error so the normal key/model retry path handles it.
  • websocket_handshake_timeout: Responses WebSocket handshake timeout for providers using that transport, mainly preset: codex with responses_websocket enabled.

These settings are provider-scoped, so project-level .chord/config.yaml can override them for one provider without changing other providers. They do not change fixed low-level connection defaults such as dial or TLS handshake timeouts.

Use max_output_tokens to set a global cap on requested output tokens. It defaults to 64000. The effective request limit is still clamped by each model’s limit.output and available total context (limit.context when known), so runtime uses the smallest applicable value across all providers.

Responses providers keep the stable Responses wire shape and do not send a max_output_tokens field on the HTTP or WebSocket request by default. For a compatible non-Codex gateway that needs an explicit server-side output cap, set compat.responses.send_max_output_tokens: true; the other Responses field toggles are available under the same provider-level object. The global value still affects Chord-side budgeting and compatibility checks when it is omitted from the wire request.

limit.input is separate: use it only for models whose providers publish an extra input cap beyond the total context window. Lowering max_output_tokens can reduce cost and long-response failure risk, but it does not increase a provider’s input allowance or replace limit.input.

max_output_tokens: 64000

Use stream_retry_rounds to put a hard ceiling on public LLM retry rounds. Each round can still walk the current model pool and each provider key in the normal order; this setting limits how many full rounds CompleteStream will make before giving up.

A “round” here means the whole public retry pass, not a single provider/model attempt. For example, stream_retry_rounds: 2 allows at most two full passes through the active routing chain. Once the cap is reached, Chord stops even for retry classes that would normally wait and continue, such as all-keys-cooling, concurrent-request 429 responses, or retryable HTTP 400 responses from a non-official compatible gateway.

Provider HTTP 400 handling is intentionally conservative:

  • Official APIs treat 400 as a terminal invalid-request error.

  • Non-official compatible gateways may return 400 for transient gateway states such as concurrency limits or upstream capacity. Those non request-shaped 400s can cool the current key, rotate to the next key, and continue after all keys are cooling.

  • Request/parameter/model-incompatible 400s still stop instead of retrying forever, for example a structured code/type such as invalid_request_error, invalid_request, or missing_required_parameter, or message-only inputs like missing required parameter, Store must be set to false, or Stream must be set to true.

  • 0 keeps the default behavior: retry until success, cancellation, or a terminal failure.

  • Positive values stop after that many rounds, even for cooling / concurrent-request retry classes.

  • This is mainly useful for automation or headless environments that prefer bounded latency over maximum persistence.

stream_retry_rounds: 3

These options affect the local TUI. They can be set in the global config and can also be overridden by project-level .chord/config.yaml when appropriate.

desktop_notification: true
desktop_notification_foreground: true
ime_switch_target: com.apple.keylayout.ABC
prevent_sleep: true
  • desktop_notification: enables terminal notifications in local TUI mode, regardless of whether the terminal is focused. Each notification pairs the terminal notification escape sequence (auto-selected by terminal, OSC 9 or OSC 777) with a terminal bell (BEL), so it can be heard even where the terminal hides notification banners while focused.

    Chord notifies when the agent actually ran and then stopped (a completed, cancelled, or loop-finished turn, or all SubAgents finishing) and for permission confirmations and questions, Handoff, loop decisions, and notify-protocol corrections waiting for input; user-initiated navigation that settles into idle (session / model-pool / MCP switches, idle slash commands) stays silent. Whether the bell is audible depends on terminal setup; see Platforms.

  • desktop_notification_foreground: controls whether notifications (both the escape sequence and the bell) are sent while the TUI is focused. Defaults to true; set it to false to notify only when the terminal is unfocused.

  • ime_switch_target: uses im-select (im-select.exe on Windows) to switch to the specified input method when entering Normal mode, and restore the previous input method when returning to Insert mode. This is useful when you want command keys to use an English keyboard layout.

  • prevent_sleep: prevents macOS idle sleep while any agent is active. It is only effective in local TUI mode.

web_fetch uses a built-in browser-like User-Agent by default. You can override it in config when a site needs a different header:

web_fetch:
user_agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/136.0.0.0 Safari/537.36

This setting works in both global config and project-level .chord/config.yaml; project config overrides the global value.

You can also configure a proxy for WebFetch requests:

web_fetch:
proxy: socks5://127.0.0.1:1080 # http, https, socks5 supported
  • proxy: nil (default): inherits the global proxy setting
  • proxy: "" (empty string): explicitly disables proxy (“direct” mode)
  • proxy: "http://...", "https://...", "socks5://...": uses specified proxy

web_fetch intentionally remains a lightweight static HTTP reader. It does not run a local browser; JS-heavy pages may be marked as Content-Quality: suspect-shell when the returned HTML looks like an application shell rather than readable content.

web_search searches through the provider’s hosted search tool and returns a summary with numbered sources. It is off by default: enable it for an Anthropic Messages or OpenAI Responses provider (or one model) whose endpoint supports hosted search.

providers:
anthropic:
type: messages
compat:
hosted_tools: [web_search]

On an OpenAI Responses provider, enable it the same way on the responses wire type:

providers:
openai:
type: responses
compat:
hosted_tools: [web_search]

compat.hosted_tools lists the hosted tools a provider’s models may serve; an omitted model field inherits the provider list, an explicit hosted_tools: [] disables all hosted tools for that model, and a non-empty list replaces the provider list. web_search is a built-in entry, declared as web_search_20250305 on Anthropic Messages and web_search on OpenAI Responses.

The tool joins the model’s tool list while its routing source contains an enabled target that can carry the declaration. Without model_pool, that source is the calling agent’s active model pool; with a named model_pool, it is the configured pool for the hosted request. Otherwise Chord withholds the tool. Each call sends a separate request carrying only the query and declares the hosted search tool there, so the main conversation request never declares it and its history stays free of provider-specific blocks. Chord returns the native results as an ordinary tool result; allowed_domains and blocked_domains travel as request parameters rather than query text.

OpenAI’s search restrictions depend on the model and the sub-request’s reasoning settings: gpt-5 with reasoning.effort: minimal does not support web search, while gpt-5.4 with reasoning.effort: none may produce lower-quality results. See OpenAI’s web search guide for per-model support.

Chord’s Responses search sub-request omits the model’s configured reasoning.effort, so it uses the server’s default reasoning settings. request_overrides can inject a reasoning parameter into that sub-request; the restrictions above apply to the settings actually sent.

The sub-request bills as tokens on the model that serves it. Providers may charge a per-search fee on top; Chord’s cost accounting counts tokens only.

For complete GPT and Claude provider recipes, see Model configuration recipes.

The top-level hosted_tools section defines provider-side (hosted) tools. Each entry becomes a local tool whose calls run one sub-request declaring the hosted tool, so a tool the provider executes server-side — web search, code execution, file search — needs configuration instead of code. A local tool appears while some provider or model enables it through compat.hosted_tools and the target’s wire type has a declaration in the entry; otherwise Chord withholds it.

Execution happens on the endpoint that serves each sub-request, so availability follows the channel rather than the model alone: a relay can serve the same models without carrying their hosted tools, and some hosted tools may be available only through the provider’s official channels. A declaration the endpoint rejects or silently drops fails the call as described for compat.hosted_tools.

hosted_tools:
code_execution:
description: Run Python code in the provider's sandbox and report its output.
parameters:
type: object
properties:
code:
type: string
description: Python source to run.
required: [code]
prompt: "Run this code and report the result:\n{code}"
read_only: true
timeout_s: 300
model_pool: tools
declarations:
messages:
tool: {type: code_execution_20250825, name: code_execution}
force: {type: tool, name: code_execution}
Field Default Description
description (empty) Local tool description shown to the model.
parameters object schema without arguments JSON Schema of the local tool’s arguments.
prompt arguments as JSON Sub-request instruction; {arg} is replaced with the matching argument value.
read_only false Marks the tool read-only for scheduling; the role permission rules still decide access.
concurrency_safe false Allows the call to run alongside other concurrent-safe read-only tools in one batch.
retry_safe false Allow retries when execution outcome is unknown; enable only when repeated execution is safe. Built-in web_search defaults to true.
image_paths (empty) Dot-separated object paths to base64 images in each call result, such as [output.image]. Images become tool result attachments; array traversal is unsupported.
timeout_s 120 Whole-call budget in seconds, covering every sub-request attempt.
declarations.<type> Wire declaration for one provider type: messages or responses.
declarations.<type>.tool (required) Wire JSON with a non-empty type, inserted into that family’s tools array. Other fields follow the provider’s schema.
declarations.<type>.force (omitted) Raw tool_choice value that forces the call. Omitted, the tool is still declared and the prompt asks for it, but a model that answers without calling it fails the call.
declarations.<type>.include (omitted) include selectors, for example [web_search_call.action.sources] on Responses.
declarations.<type>.headers (omitted) Declaration-specific HTTP headers, applied after provider headers and request overrides. Beta headers can be replaced; authentication, transport, and session headers are protected.
model_pool (empty) Named model_pools entry that serves this tool’s sub-requests instead of the caller’s own pool. Empty follows the caller: the main agent uses the main pool and a subagent uses its own. The pool must exist at startup and the calling agent must include it in its own model_pools; compat.hosted_tools still decides which entries in the pool can serve the tool. To pin a single model, create a pool containing just it — pool entries accept provider/model@variant.

Inside a declaration, {"$arg": "<name>"} is replaced with the local argument of that name; a missing or empty argument drops the key, and an object that loses all its keys is dropped too. The built-in web_search entry uses this to pass allowed_domains and blocked_domains as declaration parameters.

An entry that only exists in your configuration starts with conservative traits: not read-only, not concurrency-safe, and no automatic replay of operations with unknown outcomes. Built-in entries (currently web_search) are merged field by field, so a declaration can be retargeted to another tool version without restating the local tool surface.

Names must not collide with registered tools, reserved built-in names, or the mcp_ prefix. Startup validates the merged catalog: each entry needs an object parameter schema, a non-negative timeout, and at least one messages or responses declaration with a non-empty tool type. Headers must have valid names and values, and cannot replace authentication, transport, or session headers. Provider-specific declaration fields remain subject to the endpoint’s validation. The built-in web_search entry supports field-level overrides.

The first sub-request walks capable targets of the tool’s routing source. Unset model_pool follows the calling agent: the main agent walks its active model-pool cursor and a subagent walks its own pool. Each hosted tool remembers its successful target per routing source — shared across agents for a named model_pool, isolated per caller otherwise — and later calls start there while the pool contents stay the same; a pool rebuild or a caller pool switch starts fresh, and targets removed from the pool or no longer capable are not reused. If no target in the routing source can carry the tool, the call fails explicitly and never falls back to the main conversation’s model. For Messages pause_turn, Chord retains the complete ordered native content and sandbox container and continues on the same target, with at most four continuations under the original timeout_s budget. Token usage for all attempts belongs to the calling Agent and turn. An explicit tool_choice rejection gets one retry on the same target without forcing the call.

A named model_pool that is missing or empty prevents startup. Chord skips model entries that cannot be loaded and writes the reason to the logs. If none of the routing pool’s usable models supports the tool’s protocol and enables it through compat.hosted_tools, or the calling agent is not authorized to use the named pool, Chord hides the tool and logs the reason once for that routing source. Check the pool definition, the agent’s model_pools, and the targets’ compat.hosted_tools when a tool is missing.

With the default retry_safe: false, a possibly executed operation whose outcome is unknown (connection interruption, stream error, or exhausted continuation limit) stops automatic key/model replay. A clear rejection before execution may still try another target. Set retry_safe: true only when repeating the operation is safe.

Hosted requests share orchestration.max_active_llm_requests, provider_max_active_requests, and model_max_active_requests with other Agent requests. For a provider that allows only one concurrent request, set its entry under provider_max_active_requests to 1. Limits are local to this Chord process; separate provider entries or processes sharing one account do not share a quota gate.

For retry_safe: true tools, including the built-in web_search, transient rate limits, upstream unavailability, and transport failures get at most three rounds per target, with configured key rotation, provider backoff, and Retry-After pacing. Each round can try multiple keys; the total timeout_s budget also includes queueing and retry waits. Capacity is released while waiting between attempts. Exhausted account quota, rejected declarations, and responses without an observed hosted call do not trigger another retry round; another capable target may still be tried. Failure messages distinguish these cases and suggest an appropriate next action.

Responses remote MCP errors fail the call. Requests requiring provider-side approval stop and direct you to the local MCP integration for interactive approval; the bridge never automatically approves them. Native message citations, file references, and unknown output fields are retained. Full native output and truncated call payloads are saved as session artifacts with readable references in the tool result. Provider files currently retain container_id, file_id, and filename references; Chord does not automatically download these files. Configure image_paths to attach base64 image data actually returned by the provider.

The bridge currently supports messages and responses sub-requests. Main conversation history uses ordinary tool results. Hosted declarations for Gemini and Chat Completions, and native main-conversation replay, require separate protocol adapters.

The top-level memory section controls automatic cross-session memory extraction. Reading an existing project MEMORY.md is always automatic and needs no config; this key only decides whether Chord sends frozen history sessions to the model to grow memory records and writes project files. What gets stored, how the summary loads, and how to review or remove entries: Project Memory.

memory:
enabled: true
model_pool: memory-extract
Field Default Description
enabled false Enable automatic memory extraction for this machine + project. When on, frozen sessions may be sent to the model and auto-written into MEMORY.md / .chord/memory/records/ as ordinary project files. When off (or unset), Chord never sends history to the model and never writes memory files, but still loads an existing MEMORY.md.
model_pool (unset) Name of a model_pools entry used for extraction requests instead of the main model pool. Use it to choose extraction models and reasoning settings independently of the main conversation. When unset, extraction uses the main model pool. Both paths preserve each model’s configured reasoning settings; an omitted effort is not overridden and retains the provider default. A configured pool must be defined in model_pools; otherwise extraction stops with a setup failure naming the missing pool.
  • May appear in the global config and in project .chord/config.yaml; project values override the user-level value like every other setting.
  • Because a project can enable extraction for itself, opening a project with memory.enabled: true may start uploading that project’s history sessions to the model. When enabled, the status bar shows a MEMORY indicator so the state is visible; a stalled extraction turns it into MEMORY-FAIL and adds a one-line notice with the reason.
  • The value is read at startup; changing the config file requires a restart of the running process.

The top-level orchestration section bounds process-local resources used by MainAgent/SubAgent workflows. It does not grant tool permissions or change per-agent delegation limits such as delegation.max_children; it limits how many admitted runtimes and LLM requests can run at once, how much SubAgent input can queue, and how much mailbox data remains in memory.

Most users should keep the built-in defaults. Configure these limits when a provider has a strict concurrency quota, the host has limited memory, or orchestration metrics show sustained queueing or rejection.

orchestration:
max_live_runtimes: 10
max_borrowed_runtimes: 1
max_bypass_runtimes: 4
max_active_llm_requests: 10
provider_max_active_requests:
openai: 6
anthropic: 4
model_max_active_requests:
openai/gpt-5.5: 3
subagent_queue_messages: 256
subagent_queue_bytes: 4194304 # 4 MiB
mailbox_memory_messages: 512
mailbox_memory_bytes: 8388608 # 8 MiB
subagent_compact_usage: 0.8
waiting_main_expiry_turns: 5
waiting_main_min_wait_sec: 300 # 5 minutes
waiting_main_max_wait_sec: 3600 # 1 hour
Field Default Description
max_live_runtimes 10 Maximum normally admitted Agent runtimes. Further normal runtime acquisition waits until a slot is released. Wake reactivation may use the separately bounded borrowed or bypass pools when ordinary capacity is exhausted.
max_borrowed_runtimes 1 Additional temporary runtime admissions used to wake orchestration work that must make progress, such as a parent resuming after a child event. Borrowing is bounded separately from normal runtime slots.
max_bypass_runtimes 4 Maximum wake reactivations that may bypass both the normal and borrowed runtime pools when neither can make progress. When it is exhausted, the wake is refused and the durable message remains queued.
max_active_llm_requests 10 Process-wide maximum concurrent LLM requests across orchestrated agents. Eligible requests wait when the limit is full.
provider_max_active_requests none Optional concurrent-request limits keyed by provider name, for example openai. A request must satisfy this limit and the process-wide limit.
model_max_active_requests none Optional concurrent-request limits keyed by provider/model. Inline variants such as @high are ignored for matching, so openai/gpt-5.5 covers all variants of that model.
subagent_queue_messages 256 Maximum pending input messages for each SubAgent. A new enqueue is rejected when either this count or the byte limit is reached; existing queued messages are preserved.
subagent_queue_bytes 4194304 Maximum estimated bytes of pending input for each SubAgent. This is an in-memory admission bound, not a disk spool.
mailbox_memory_messages 512 Maximum SubAgent mailbox messages retained in memory across the MainAgent inbox and owner-specific mailboxes.
mailbox_memory_bytes 8388608 Maximum estimated bytes retained by those in-memory mailboxes. Durable non-progress messages that exceed the memory budget are referenced through the on-disk mailbox spool; progress updates may be coalesced or omitted from memory.
subagent_compact_usage 0.8 Proactively compress a SubAgent’s context when estimated usage reaches this fraction of its usable input budget. The default matches context.compaction.threshold; SubAgents use local token estimates and a lightweight sliding-window checkpoint rather than MainAgent’s usage-driven compaction pipeline. Must be greater than 0 and less than 1.
waiting_main_expiry_turns 5 User-turn budget for a SubAgent parked while waiting for its owner. The turn budget expires only after waiting_main_min_wait_sec has also elapsed; waiting_main_max_wait_sec still expires the wait unconditionally.
waiting_main_min_wait_sec 300 Minimum wall-clock wait, in seconds, before the turn budget can expire a waiting_main task.
waiting_main_max_wait_sec 3600 Maximum wall-clock wait, in seconds, after which a waiting_main task expires regardless of user-turn activity. The effective value is never below waiting_main_min_wait_sec; when both clocks are set explicitly and this maximum is below the minimum, loading fails instead of clamping.
  • These settings may appear in the global config and in project .chord/config.yaml. Positive project scalar values override the corresponding global values.
  • provider_max_active_requests and model_max_active_requests are merged by key. A project entry replaces the same global key while preserving unrelated global entries.
  • Scalar values that are zero or negative do not mean “unlimited”: they retain the inherited or built-in default. subagent_compact_usage is only valid strictly between 0 and 1: an out-of-range value (including 0) is ignored with a warning, a project value then inherits the merged global value, and an unset global falls back to 0.8. Unlike context.compaction.threshold: 0, zero does not disable SubAgent context protection.
  • Only positive provider/model map limits are enforced. Keep map keys explicit and use positive integers; do not rely on zero as a general unlimited-mode switch.
  • Limits are process-local. They do not coordinate quotas across multiple Chord processes.
  • To comply with an API quota, set the provider or model limit first; keep max_active_llm_requests as the overall safety ceiling.
  • Keep max_bypass_runtimes small and positive. It exists only to let wake reactivations make progress when normal and borrowed capacity are exhausted; it is not ordinary throughput capacity.
  • On a memory-constrained host, reduce mailbox byte/message limits gradually. Overflow uses durable storage, so lower limits trade memory for additional disk I/O.
  • Reduce SubAgent queue limits only when producers can handle enqueue rejection. These queues do not spill to disk, and overly small limits can interrupt parent/child coordination.
  • Keep max_borrowed_runtimes small but positive. Borrowed slots exist to break orchestration progress stalls, not to increase ordinary throughput.
  • A waiting_main task expires when its turn budget and minimum wait are both satisfied, or when the maximum wait is reached. Increase the turn budget or minimum wait when owners need more time to respond; increase the maximum only when parked tasks should remain recoverable for longer.
  • Lowering subagent_compact_usage reduces context-overflow risk but causes earlier and more frequent compression. Raising it reduces compression work but leaves less recovery headroom.
  • Increasing concurrency is not automatically faster: provider throttling, model latency, local memory pressure, and workspace lease contention can reduce effective throughput. Change limits using observed queue/rejection metrics and end-to-end latency rather than CPU count alone.

MCP servers connect in two ways: Chord launches a local command and exchanges JSON-RPC over stdio, or it connects to a remote HTTP endpoint.

mcp:
chrome-devtools:
command: "npx"
args: ["-y", "chrome-devtools-mcp@latest"]

command is the executable to launch, args are its arguments, and optional env entries are appended to the inherited environment. Chord starts the process and talks to it over stdin/stdout using newline-delimited JSON-RPC.

MCP servers can expose many tools. Use allowed_tools to expose only selected remote tool names and avoid sending unused tool schemas to the model:

mcp:
search:
url: https://mcp.exa.ai/mcp
allowed_tools:
- web_search_exa
- web_fetch_exa

The server name (search above) is user-defined. With this example, Chord registers only mcp_search_web_search_exa and mcp_search_web_fetch_exa. Filtered tools are not registered and do not enter the LLM tool surface.

For remote MCP servers that require authentication, use headers to send extra headers with every request. A service like Exa requires an x-api-key:

mcp:
exa:
url: https://mcp.exa.ai/mcp
headers:
x-api-key: "$EXA_API_KEY"

A header value starting with $ is expanded from the environment (here EXA_API_KEY), so secrets do not have to be written into the config file; a $ value that expands to an empty string is a configuration error, since it would authenticate with a blank credential. Header names must be valid HTTP header names, and values must not contain CR or LF.

headers applies only to remote (url) servers; stdio servers do not carry HTTP requests, and configuring headers for one is rejected. Protocol-managed headers (Content-Type, Accept, Mcp-Session-Id) are set by Chord and are not affected by headers.

By default, configured MCP servers auto-start and become part of the default LLM tool context. For an MCP server you do not need in every conversation, set manual: true: it stays disabled at startup, Chord normally does not connect to it, and its tool descriptions are not added to the default context, reducing context overhead. Enable it manually only when you need it:

mcp:
exa:
url: https://mcp.exa.ai/mcp
manual: true
  • When manual: true, the server starts in a disabled (gray) state and does not connect until you enable it.
  • Only servers configured with manual: true can be changed at runtime with /mcp. Auto-start servers are read-only in the MCP selector and are not affected by /mcp enable|disable.
  • Enable/disable at runtime with /mcp (menu in TUI) or with explicit commands:
    • /mcp enable <server>
    • /mcp disable <server>
    • /mcp status
  • Runtime /mcp enable|disable changes are allowed while a turn is running. The current in-flight request keeps the tool surface it started with; the next LLM request (including automatic retry/recovery requests) applies the new execution state.
  • By default, that next request rebuilds the top-level MCP tool surface and may miss the existing prompt cache. Models that explicitly enable compat.chat_completions.mcp_system_tools_message or compat.responses.mcp_additional_tools instead receive request-only tool declarations at a fixed conversation anchor. Later requests replay each declaration at the same position; disabling a server blocks execution but keeps its existing declaration, preserving the prefix. The mount follows the selected model: a model switch rebuilds the top-level tool surface for the new model and then resumes dynamic mounts where that model accepts them. Session resume or switch and durable compaction instead revert to top-level tools for the rest of the session run, because the fixed anchors cannot be trusted; a later model switch keeps that revert, and only starting a new session run (such as /new) re-enables dynamic mounts. A tool whose earlier calls are still in the history but that has no declaration in the current run also makes the request fall back to top-level tools, so its declaration never lands after its own calls. The revert itself is silent; a later /mcp enable|disable warns about the prompt cache again only once a request has installed the rebuilt surface.
  • The enabled/disabled intent for manual servers is saved with the session: /mcp enable persists it, /mcp disable clears it, and resuming the session (including after restart) reconnects the manual servers that were enabled when it was last active. A connection failure keeps the intent, so the server stays “enabled (unavailable)” and can be retried instead of silently returning to disabled.

Auto-start MCP servers still connect asynchronously after the TUI starts, but the first LLM request waits until each auto-start server has either connected successfully or reached a terminal failure state. This avoids tool-surface inconsistency between the agent and the model.

Built-in roles include builder and planner. Both are main-mode, so delegate is not registered until you define at least one mode: subagent role of your own; see mode below. You can also add custom agents or override built-ins. Agent files can live in:

  • ~/.config/chord/agents/
  • .chord/agents/

Supported file formats:

  • .md: YAML frontmatter plus a Markdown body. The body becomes the system prompt.
  • .yaml / .yml: plain YAML. Use prompt or system_prompt for the system prompt.

Markdown agent example:

---
name: backend-coder
description: Backend developer
mode: subagent
permission:
write: ask
edit: ask
---
You are an agent focused on backend development.

Equivalent YAML agent example:

name: backend-coder
description: Backend developer
mode: subagent
permission:
write: ask
edit: ask
prompt: |
You are an agent focused on backend development.

Common fields include:

  • name: agent name. If omitted, Chord uses the filename without extension. If specified, it must match the filename without extension (for example, builder.yaml must declare name: builder). A single directory cannot contain duplicate agent names, including duplicates across .md, .yaml, and .yml. Project-level agents may still override same-named global agents by design.
  • description: short description shown to the main agent when delegation is available. Put routing intent here; Chord does not keep a separate annotation layer for preferred tasks or write mode. Whether a role may write files is decided by permission, and the Delegate tool surfaces that as empty_scope=allowed or non_empty_scope=required on each agent choice.
  • mode: main for a MainAgent role, or subagent for a SubAgent. Empty and unknown values behave as main; sub_agent and sub are accepted as SubAgent aliases. The delegate tool is registered only when at least one subagent role is visible to the delegating role, so a configuration with no subagent definitions has no delegation surface at all: that is the usual reason delegate appears to be missing.
  • model_pools: optional ordered list of pool names this agent can use. Pool definitions live in config.yaml top-level model_pools; when omitted, the agent can use all top-level pools sorted by name. Inline variants such as openai/gpt-5.5@high are specified in the pool definitions.
  • variant: default variant when a model ref does not include @variant.
  • permission: per-tool permission policy for this agent. Permissions live directly in agent config files; when the confirmation popup remembers a rule, project updates the current project’s .chord/agents/<role>.yaml, and global updates the user config directory’s agents/<role>.yaml (default: ~/.config/chord/agents/<role>.yaml). Chord no longer writes a separate permissions directory. Some orchestration tools have special semantics (delegate patterns match agent_type and also gate delegated-work controls such as cancel; handoff and done treat allow and ask as workflow-available states with Chord’s own confirmation gates). See Permissions & Safety before relying on fine-grained control-tool rules.
  • mcp: additional auto-start MCP servers scoped to this agent. Agent MCP is additive: a server name already present in the effective global/project mcp config is a startup error. Agent-scoped servers cannot use manual: true because runtime MCP controls manage the top-level server surface; configure a manual server at the project/global level instead. Remove an agent entry to inherit a top-level server, rename it for a separate private server, or override the top-level server in .chord/config.yaml for the whole project. Different agents may reuse the same private server name without sharing the connection unless they are instances of the same agent definition.
  • delegation: delegation limits for this agent definition. Values above the ceiling or negative values are configuration errors:
    • max_children: how many direct, still-active child tasks this agent may have at once. Defaults to 10; the ceiling is 64.
    • max_depth: how deep nested delegation may go. It is evaluated per worker, against the worker’s own definition: a SubAgent’s ability to delegate further is checked against its own delegation.max_depth and its current depth, never against its parent’s or the root role’s setting, so a root role with max_depth: 1 cannot stop a child definition that declares max_depth: 8 from nesting deeper. Defaults to 1 (a first-level SubAgent cannot delegate further until its own definition raises the value); the ceiling is 8.
    • child_join: whether children a SubAgent delegates stay tied to the owner’s task. Defaults to true: the owner cannot complete while joined children are still running, so its completion is deferred until they finish or are explicitly stopped, and a cancelled or failed owner cancels its joined children with it. With false, the owner may finish early and its still-running children detach and continue under the main agent instead of being cancelled. Only nested delegation is affected: children delegated by the main agent never join, because the main agent is not itself a task.
  • prompt / system_prompt: system prompt for plain YAML files. Setting either one replaces any built-in prompt block the role would otherwise get.
  • prompt_preset: selects a built-in role prompt block by capability instead of by role name. Accepted values are planning and none. planning injects the built-in planning block (plan-document naming and format, the direct-answer-versus-plan decision, handoff ordering, and plan quality rules); it also suppresses the bug-triage block, which would otherwise duplicate the planning workflow’s own investigation outline. none suppresses any built-in block. When the field is omitted, the role gets no built-in block whatever it is called: the role name never selects one, so a custom role named planner has to declare prompt_preset: planning to keep the planning block. Unknown values are a configuration error.
  • prompt_append: text appended after the effective role prompt: after the preset block, or after prompt / system_prompt when the role replaces it. Use this to add project conventions without taking over maintenance of the whole block, which also keeps the preset’s tool-aware wording (it adapts to the tools the role can actually see).

A custom planning role that reuses the built-in block:

name: architect
description: Architecture planning role
mode: main
prompt_preset: planning
prompt_append: |
Reference the relevant ADR number in every plan document.
permission:
"*": deny
read: allow
grep: allow
glob: allow
write:
.chord/plans/*: allow
handoff: allow

Example:

name: builder
mode: main
model_pools: [default]
permission:
"*": deny
read: allow
view_image: allow
grep: allow
glob: allow
web_fetch:
"localhost:8000": ask
shell: allow
edit: ask
write: ask

In an allowlist like this, the leading "*": deny covers every tool you did not list, so anything the role needs must be named. Two exceptions are worth knowing, because they would otherwise look like the feature is broken:

  • compact_context and done are not covered by the wildcard. The switch that makes each reachable (context.compaction.model_driven for the former, starting a loop for the latter) is itself the authorization, so this role can run model-driven compaction and loop mode without listing them. Name a tool explicitly (done: deny) when you do want to withhold it. See Permissions & Safety.
  • Everything else is covered normally. todo_write and question are ordinary tools here: leaving them out means the model tracks no TODO list and asks the user in plain assistant text instead of a structured prompt. Both degrade cleanly, so add them only if you want those channels.

Long-session context handling covers context compaction (LLM-generated summaries that rewrite session history) and context reduction (request-time trimming of stale tool output). Both are configured under the top-level context: key and documented on their own page: Context management.

After edit, apply_patch, or write modifies a file, Chord can append language diagnostics to the tool result so the model sees compile or lint problems immediately. This is controlled by the diagnostics config and is enabled by default for Python (an LSP semantic backend with a Ruff quick fallback). Set diagnostics.enabled: false to skip the whole pipeline.

Native file tools also send workspace/didChangeWatchedFiles events to matching LSP servers before syncing the textDocument: write sends Created for new files, write on existing files plus edit / apply_patch send Changed, and successful delete sends Deleted. This helps Pyright, TypeScript, gopls, rust-analyzer, and similar servers refresh their project graph promptly, reducing transient unresolved-import/module diagnostics after new files are created.

Diagnostics are still returned immediately in file-tool results so the model can attribute problems to the current edit; files created or removed by shell commands or external programs are not yet reported through a full filesystem watcher.

For Python, two backends are used:

  • diagnostics.python.semantic_backend: the primary LSP server (default pyright). Its server field must match a server key under lsp so the language server is actually configured.
  • diagnostics.python.quick_backend: a one-shot fallback (default ruff check) used for large files, or when the semantic backend is unavailable.

diagnostics.python.large_file.{line_threshold, byte_threshold, strategy} decides when a file is large enough to use the quick backend instead of the semantic one; run_semantic_when_quick_unavailable: true forces the semantic backend even on large files when the quick backend is missing. Ruff quick diagnostics do not update the LSP sidebar: they appear only in edit, apply_patch, or write results and note that full semantic diagnostics were skipped.

Recommended Python skeleton:

lsp:
pyright:
command: pyright-langserver
args: ["--stdio"]
file_types: [".py", ".pyi"]
diagnostics:
python:
semantic_backend:
server: pyright
quick_backend:
type: command
command: ruff

diagnostics.python.output.{max_near_diagnostics, max_outside_diagnostics, max_total_diagnostics, near_range_before_lines, near_range_after_lines} shapes how much appended diagnostics text is shown, prioritizing errors and warnings before info and hints. See the Configuration cheatsheet for the full field list.

Diagnostics appended to the tool result cover the edited files’ own problems, plus cached problems from other files in the same directory as an edited file (Go packages are compiled per directory, and workspace diagnostics cover every file in a diagnosed package). Each other-file diagnostic is attached only once per session: the same problem is not repeated in later tool results until that diagnostic disappears from the server’s published set, after which a reappearing problem is reported again.

Later edits do not repeat it either: a problem the model already has costs context to restate, so a surviving diagnostic stays suppressed and only changes are reported. A problem that was fixed (by another agent, a file copy, or a git checkout restore) stops being reported instead of being served from cache: the cached diagnostics are withheld as soon as the file no longer matches what they were computed from, and they are dropped once the server publishes without them.

Resuming a session keeps the suppression instead of restarting it: diagnostics already rendered in the restored transcript are recovered from it, so --continue does not re-announce problems that are already visible earlier in the same conversation.

Diagnostics for a file that changed on disk since the server last published them (for example, fixed by another editor or process) are skipped until Chord synchronizes the file and receives fresh diagnostics, because the cached result may no longer reflect its current content.

Terminal window
# smoke-test all providers with representative models
chord doctor models
# test one provider's representative model
chord doctor models --provider openai
# test an exact model or variant
chord doctor models --model openai/gpt-5.5@high
chord doctor models --provider openai --model gpt-5.5@high
# audit each entry in a model pool independently
chord doctor models --pool thinking

Use this command as an auth, endpoint, transport, model, and variant tuning smoke test. It uses the same merged global + project config view as normal runtime startup, so project-level provider/proxy/model overrides are included. Pool diagnostics request each pool entry independently rather than following the normal fallback chain.

The full top-level keys of config.yaml (both global ~/.config/chord/config.yaml and project-level .chord/config.yaml). All keys are optional unless noted.

Key Type Default Scope Summary
providers map[name]Provider — global / project Per-provider config (type, api_url, preset, key_rotation, key_order, models, compress). See Minimal provider config.
model_templates map[name]YAML empty global / project YAML-anchor namespace only; entries are reusable through aliases and are not runtime model definitions.
model_pools map[name][]ref — global / project Reusable named pools of full provider/model[@variant] refs. See Model pools.
thinking_translation object disabled (max_chars: 1000) global / project Optional appended translation preview for thinking / reasoning cards. Requires target_language and model_pool; failures only skip the affected thinking block.
context object see below global / project compaction and reduction settings. See Context compaction and Context reduction.
diagnostics object enabled (Python LSP + Ruff fallback) global / project Post-tool diagnostics appended to edit, apply_patch, or write results. diagnostics.python.semantic_backend is the primary LSP server (default pyright); diagnostics.python.quick_backend is a one-shot fallback (default ruff check). diagnostics.python.large_file.{line_threshold, byte_threshold, strategy} controls when large files use the quick backend, and run_semantic_when_quick_unavailable: true forces semantic diagnostics when the quick backend is missing. diagnostics.python.output.{max_near_diagnostics, max_outside_diagnostics, max_total_diagnostics, near_range_before_lines, near_range_after_lines} shapes the appended diagnostics text. Diagnostics are shown by severity priority (errors/warnings first, then info/hints if slots remain). Set diagnostics.enabled: false to skip the whole pipeline.
skills object empty global / project paths: [...] — additional skill directories beyond the defaults.
confirm_timeout int (seconds) 0 (no timeout) global / project Timeout for confirmation dialogs in TUI; 0 means wait forever.
question_timeout int (seconds) 0 (no timeout) global / project Timeout for the Question tool in TUI and headless; 0 means wait forever. The countdown covers the whole wait, including time queued behind another dialog, and each question in a batch is timed separately. When it elapses, the question closes as no_response; Chord never adopts an answer automatically.
diff object {inline_max_columns: 200} global / project TUI diff rendering. inline_max_columns caps one-line inline diff width.
desktop_notification bool false global / project Enable local-TUI terminal notifications; Chord auto-selects OSC 9 or OSC 777 by terminal and pairs each notification with a terminal bell (BEL). (Unsupported terminals ignore the sequence; see Platforms.)
desktop_notification_foreground bool true global / project Send local-TUI terminal notifications (escape sequence and bell) while the terminal is focused. Set to false for background-only notifications.
prevent_sleep bool false global / project Prevent macOS idle sleep while any agent is active. macOS-only; no-op elsewhere.
keymap map[action][]key see Keybindings global / project Override key bindings. Action names use lower snake_case.
commands map[/cmd]text empty global / project Custom slash commands; "/cmd" → text inserted as a user message. See Customization: Custom slash commands.
ime_switch_target string empty global / project IM identifier passed to im-select / im-select.exe when entering Normal mode. Linux/macOS/Windows.
log_level string info global / project debug / info / warn / error. debug is verbose.
paths object XDG defaults global only state_dir, cache_dir, sessions_dir, logs_dir. CLI flags and CHORD_* env vars override.
maintenance object disabled global only size_check_on_startup, warn_state_bytes, warn_cache_bytes.
lsp map[name]Server empty global / project Per-language-server config. See Customization: LSP.
mcp map[name]MCP empty global / project / agent Per-MCP-server config. See MCP.
hooks object empty global / project / agent Hooks per trigger point. See Hooks.
max_output_tokens int 64000 global / project Global cap on requested output tokens. Effective limit is also clamped by each model’s limit.output; reasoning requests also respect it.
stream_retry_rounds int 0 (retry until success/cancel) global / project Hard cap on public LLM full-round retries. 0 keeps retrying until success, cancellation, or terminal failure.
proxy string empty (use env / direct) global / project Global proxy URL. Per-tool override via web_fetch.proxy.
web_fetch object empty global / project user_agent, proxy (inherits global if nil; empty string = direct). See WebFetch.
worktree object empty global / project branch_prefix names worktree branches; root sets where checkouts are created (a relative value resolves against the repository root; the default stays outside it under the state dir). See Worktrees.

Chord automatically propagates the current Chord session id to OpenAI-family providers as cache/routing affinity metadata: OpenAI Responses requests include prompt_cache_key, and OpenAI Chat Completions / Responses HTTP requests include X-Session-Id and session-id headers when a session id is available.

The key is per client rather than per provider: the main agent uses the current Chord session id, and each SubAgent derives its own <session>:sub:<instanceID> key so one agent’s requests never inherit another’s cache identity. These fields are not user-configurable; they follow the active Chord session and are cleared or changed on session switch/resume.

Anthropic prompt caching is driven by cache_control blocks, and Chord also sends JSON-formatted metadata.user_id automatically with a stable anonymous device_id plus a stable routing session_id derived from local/provider identity. These Anthropic metadata fields are not user-configurable.

In explicit mode (the default for Anthropic models), Chord places up to four cache_control breakpoints by priority: the last system block, the frozen reduced-prefix boundary (when incremental reduction has frozen a stable prefix), the newest durable message, and the last assistant message, so long agent loops reuse the frozen historical surface instead of re-writing the moving tail each turn.

The newest breakpoint deliberately skips request-scoped overlays (runtime hints appended after the conversation tail), because those bytes are gone on the next request and a cache entry written past them could never be read back.

For Anthropic models, prompt_cache.ttl accepts 5m (the default when omitted) and 1h, and applies to every breakpoint Chord places in both auto and explicit mode:

providers:
anthropic:
models:
claude-sonnet-4-5:
prompt_cache:
ttl: 1h

Gemini does not have a simple per-request session-id cache key in Chord’s generateContent transport; its cache signals come from provider-specific cached-content APIs/usage fields, not from a Chord session id header.

Field Type Description
type string messages / chat-completions / responses / generate-content. Auto-detected from api_url or preset when omitted.
api_url string Endpoint URL. Chord detects provider type from the URL path, ignoring query strings and fragments. For Gemini, the /models base path; Chord appends /{model}:streamGenerateContent?alt=sse. For Azure Responses, ?api-version=... is optional and can be used to pin a specific API version.
preset string codex (OpenAI Codex / ChatGPT OAuth). Azure OpenAI Responses uses a plain type: responses provider with auth_scheme: api-key, store: true, and compat.request_overrides.headers set to null for the Codex identity headers.
trust_http_400 bool Treat HTTP 400 as a terminal request error. preset: codex defaults to true; aggregating/proxy gateways default to false because they often wrap upstream overload as 400.
retry_after_max_s int Longest Retry-After wait honored, in seconds (1-86400). The header always applies as the key cooldown, ahead of retry_backoff/retry_delay_ms; this only bounds how long a single hint may block a key. preset: codex defaults to 86400; third-party gateways, which can echo arbitrary values, default to 60.
key_rotation string on_failure (default) / per_request. Controls when a credential / API key is reselected.
key_order string sequential (non-Codex default) / random / smart (Codex only). Controls how Chord chooses among selectable keys.
retry_backoff string exponential (default) / fixed / none. Controls generated delay between complete rounds and, when explicitly set, ordinary HTTP 429 key cooldown. Explicit settings replace Retry-After for ordinary 429s; confirmed quota resets and hard credential states still win.
retry_delay_ms int Base/fixed round and ordinary-429 delay in milliseconds, from 0 through 60000; 0 / omitted defaults to 1000ms. Setting the field—including explicit 0—is an override even when retry_backoff is omitted. Ignored for none. Out-of-range values are logged and fall back to the default instead of failing startup.
compress string Upstream request body compression encoding: gzip or zstd; unset = off. Applies only when compression shrinks the payload. The boolean compress: true form is gone — it is ignored and reported by chord doctor config (migrate to compress: gzip).
response_header_timeout int Timeout in seconds from starting a streaming HTTP request until response headers arrive, including connection setup and request-body upload. 0 / omitted uses the built-in default; healthy streams are bounded by stream_idle_timeout, not a total request timer.
stream_idle_timeout int Stream idle timeout in seconds for this provider. 0 / omitted uses built-in SSE/WebSocket idle defaults.
stream_total_timeout int Wall-clock cap in seconds for one stream, including Codex Responses WebSocket reads. 0 / omitted applies no cap — a stream that keeps producing data is never cut off by elapsed time alone.
websocket_handshake_timeout int Responses WebSocket handshake timeout in seconds. 0 / omitted uses the built-in default.
supported_service_tiers list Provider-level default accepted non-standard tiers for its models, e.g. [fast, slow] or [fast]. Model entries can override it.
parallel_tool_calls bool true — Provider-level default for Responses / Chat Completions tool parallelism; model and variant values override it.
compat.responses.* object protocol defaults — Provider-level optional Responses fields: send_store, send_reasoning_include, send_tool_choice, send_prompt_cache_key, send_max_output_tokens, and mcp_additional_tools.
compat.responses.mcp_additional_tools bool false — Mount runtime manual-MCP schemas as fixed-anchor input[type="additional_tools"] items instead of changing top-level tools. Enable only for Responses endpoints/models known to accept this item. The mount follows the currently selected target; when a fallback pool member without this capability serves the request, Chord inlines the declarations into that request’s top-level tools array.
compat.hosted_tools list (empty) — Hosted tool names this provider’s models may serve. Each call runs as a separate request that declares the tool’s per-type wire declaration and returns the provider’s result as an ordinary tool result; the main conversation request never declares it. Enable only on endpoints known to support the declaration: one that rejects it fails the call with the endpoint’s error, and one that silently ignores it fails with a message pointing back at this list. The catalog shape is documented under Hosted tools.
compat.apply_patch.enabled bool Three-state — when omitted, Chord infers from the model name. true keeps apply_patch (hiding edit, write, and delete); false falls back to edit with write/delete visible. gpt-5-and-later family names (gpt-5, gpt-5-mini, gpt-5-nano, gpt-5-codex, any gpt-5.* name, later majors like gpt-6-astra) and codex-auto-review default to true; gpt-oss-*, gpt-3.5, gpt-4/4o, the o-series, and non-OpenAI models default to false.
compat.apply_patch.freeform bool Three-state — when omitted, Chord infers from the model name and the wire type. true emits apply_patch as a freeform custom tool (type: "custom" with a grammar); false emits a JSON function tool. gpt-5-and-later family names and codex-auto-review on Responses endpoints default to true; all non-Responses wires default to false (they have no custom tool type). Hosts that accept Responses but reject custom tools have no built-in exception: set false there, or true only for gateways that actually accept custom tools.
compat.chat_completions.send_stream_options bool true — Omit stream_options.include_usage for gateways that reject it; streaming token usage then remains unavailable.
compat.chat_completions.infer_finish_reason bool false — Derive a normal stop / tool_calls completion for compatible gateways that end the stream without emitting finish_reason; otherwise those streams are treated as interrupted.
compat.chat_completions.requires_tool_result_name bool false — Emit the paired tool name on tool result messages for gateways that require name alongside tool_call_id.
compat.chat_completions.requires_assistant_after_tool_result bool false — Insert a synthetic assistant message between a tool result and the next user message for gateways that reject a user message directly after tool results.
compat.chat_completions.mcp_system_tools_message bool false — Mount runtime manual-MCP schemas as fixed-anchor role: system messages with a tools field and no content, instead of changing top-level tools. Enable only for models known to accept the Kimi-compatible dynamic-tool shape. The mount follows the currently selected target; when a fallback pool member without this capability serves the request, Chord inlines the declarations into that request’s top-level tools array.
compat.chat_completions.keep_reasoning_effort bool false — Keep reasoning_effort and reasoning request overrides active when a current-turn assistant tool-call message replays without reasoning_content. Chord otherwise reads the missing content as a backend that cannot replay reasoning and strips those controls for the rest of the turn; enable it for endpoints that accept reasoning controls without a reasoning-content replay contract, such as Grok on the Chat Completions wire. It only keeps the request-side controls in place; it does not supply reasoning content to backends whose contract validates the replayed history (DeepSeek when a request carries tools, Kimi K3, Qwen preserve_thinking).
compat.chat_completions.native_thinking string Selects the request shape Chord uses for a model’s thinking settings when the endpoint is a Chat Completions gateway that translates the call into the model’s native API. Without a value, only a DeepSeek route selects the thinking:{type} shape: a model ID that names a DeepSeek API model, or any model under compat.reasoning_continuity.contract: deepseek (contract: none drops the name shortcut). Every other model needs an explicit value. Values: gemini (extra_body.google.thinking_config), gemini-3 (the same shape with Gemini 3 signature repair), anthropic (thinking:{type,budget_tokens}), thinking (the native thinking:{type} object used by DeepSeek, GLM, Kimi K2.x, and Doubao), and qwen (enable_thinking); family names such as claude, deepseek, glm, kimi, and doubao are accepted as aliases. off (alias none) disables the conversion for endpoints that reject unknown body fields. A model that configures no thinking block never sends the field, except on a DeepSeek route: its reasoning contract enables thinking by default, while explicit thinking.type: disabled or reasoning.effort: none disables it and omits effort (see compat.reasoning_continuity.contract). The selector is also the only signal that a gateway model is Gemini or Claude: without it Chord does not write Gemini thought signatures back, and Gemini 3 rejects the request that follows each tool call (HTTP 400), so pin gemini-3 on every Gemini 3 model behind a gateway. chord doctor config warns about such models. An explicit selector governs this request shape only and always wins over the DeepSeek default; the reasoning contract stays in force, and off does not disable it. See Thinking behind a Chat Completions gateway.
compat.usage.input_includes_cache_read bool Protocol default — Override whether the provider’s top-level input count already contains cache-read tokens. Defaults: Messages false; Chat Completions / Responses / Generate Content true.
compat.usage.input_includes_cache_write bool Protocol default — Override whether the provider’s top-level input count already contains cache-write/cache-creation tokens. Defaults: Chat Completions / Responses true; Messages / Generate Content false.
models map Map of model id → model config.
Field Type Description
limit.context int Total request window in tokens when known. If limit.input is omitted, Chord derives the input budget from this minus the model’s limit.output (falling back to the max_output_tokens default when no output cap is declared).
limit.input int Separate input cap when a provider publishes one. Chord uses it to compact or retry before the prompt is too large.
limit.output int Maximum output tokens; runtime is also clamped by max_output_tokens.
compaction object Per-model compaction overrides: compaction.threshold (auto-compaction usage ratio; 0 disables for this model) and compaction.reminder (pressure-reminder line; derived from threshold when absent, -1 disables the reminder only). Unset fields inherit the global context.compaction.*. Out-of-range values are rejected with a warning and inherit the global value. Derivation and tuning guidance: Context compaction.
reasoning object OpenAI reasoning options. reasoning.effort passes through without a local whitelist, so any provider-supported level (e.g. GLM max / minimal / none) reaches the upstream unchanged; Responses normalizes whitespace and casing before sending (unset = omit and use provider/model default). For Responses, reasoning.summary supports auto / concise / detailed / none; when reasoning is active, unset defaults to auto, while none opts out explicitly.
text.verbosity string Optional OpenAI text verbosity hint where supported; leave unset to use the provider/model default unless you intentionally want low / medium / high.
thinking object Extended-thinking options. Messages: type: adaptive carries no token budget and pairs with thinking.effort, which Chord sends as output_config.effort; Claude type: enabled requires thinking.budget; DeepSeek instead accepts type: enabled with thinking.effort and no budget; display applies only to enabled / adaptive. Gemini: thinking.level / thinking.budget / thinking.include_thoughts map into the generation request (see Google Gemini).
compat.reasoning_continuity.mode string Optional continuity override. Use openai_visible for Chat Completions models that require unchanged assistant reasoning_content and can accept portable visible reasoning from other wires; it also enables the missing-reasoning_text fallback for Responses targets with that continuity contract. Use anthropic_unsigned only for verified Messages-compatible models that replay or accept visible unsigned thinking; use none to opt out of a provider-level default. DeepSeek Chat/Messages targets ignore this field, including none: their reasoning contract fixes the mode (openai_visible on Chat, anthropic_unsigned on Messages).
compat.reasoning_continuity.contract string Declares an endpoint-specific request contract, overriding what the model name implies. deepseek selects the DeepSeek tool-history passback rules and request tuning on the Chat Completions and Messages wires (on Chat Completions also the thinking:{type} shape while native_thinking is unset), which an alias whose model ID does not identify the backend needs; gemini-3 enables missing thought-signature repair on the native Gemini endpoint (behind a Chat Completions gateway, native_thinking: gemini-3 does this instead); none opts a route out of whatever its model name implies, including that shape, for example a third-party route whose model ID names a DeepSeek API model but serves another backend, on the Chat Completions and Messages wires alike. A model ID that names a DeepSeek API model selects deepseek on its own, and on the native Gemini endpoint a model ID whose final component starts with gemini-3 selects gemini-3; other endpoints keep the generic continuity behavior. An unknown value is rejected by the config loader.
compat.reasoning_continuity.reasoning_replay string How much reasoning from completed turns is replayed. DeepSeek Chat/Messages always replay the complete history regardless of this window. current_turn (default) keeps only reasoning after the last user message; all replays completed turns unchanged for backends whose contract requires the full assistant history (Kimi K3 / keep: all, Qwen preserve_thinking, GLM clear_thinking: false, Xiaomi MiMo); none strips reasoning everywhere including the current turn, for endpoints verified to accept that.
compat.forced_tool_choice.suppress_in_thinking bool Downgrade loop-forced tool_choice: required to the backend default while reasoning/thinking is active, for OpenAI-compatible endpoints that reject forced tool choice in thinking mode.
compat.forced_tool_choice.auto_only bool Downgrade any non-auto tool_choice to the backend default unconditionally, for backends that only support tool_choice: "auto". Chat Completions / Messages / Gemini omit the field; Responses still sends "auto" unless compat.responses.send_tool_choice is false.
compat.request_overrides.body object Recursive JSON patch applied after Chord constructs the protocol request. null deletes a field.
compat.request_overrides.rename_body_fields map Renames final JSON fields while preserving Chord’s computed values. A null target deletes the source field.
compat.request_overrides.headers map Sets final request headers. A null value removes that header.
compat.chat_completions.mcp_system_tools_message bool Model-level override for the provider default described above.
compat.chat_completions.keep_reasoning_effort bool Model-level override for the provider default described above.
compat.chat_completions.native_thinking string Model-level override for the provider default described above.
compat.responses.mcp_additional_tools bool Model-level override for the provider default described above.
compat.hosted_tools list Model-level list that replaces the provider default described above.
compat.apply_patch.enabled bool Model-level override for the provider default described above.
compat.apply_patch.freeform bool Model-level override for the provider default described above.
variants map Named parameter presets. Reference with provider/model@variant.
modalities.input array Subset of text / image / pdf. Defaults to [text]; declare image / pdf explicitly when supported.
supported_service_tiers list Provider-level default or model-level override for accepted non-standard tiers, e.g. [fast, slow] or [fast]. Omit to use preset defaults.

Service tiers and prompt caching are provider-specific. OpenAI supports priority/flex-style tiering plus prompt_cache_key / prompt_cache_retention; Anthropic supports cache_control with 5m and 1h TTLs and service-tier controls; Gemini uses its own routing / thinking / cached-content mechanisms when available. Chord maps the user-facing tier to the closest supported provider behavior instead of forcing one wire format across all backends.

For a compatible gateway whose usage fields differ from its declared protocol, set the usage semantics explicitly. For example, a Responses-compatible gateway that reports uncached input separately from both cache buckets needs:

compat:
usage:
input_includes_cache_read: false
input_includes_cache_write: false

A Messages-compatible gateway that reports inclusive usage (input_tokens is the full input including cache hits, and cache_read_input_tokens is only the hit subset) needs the opposite override. Otherwise, Chord counts the cache reads twice and understates the cache-hit rate:

compat:
usage:
input_includes_cache_read: true

OpenAI reasoning items can also be returned as reasoning.encrypted_content when you need stateless continuation. Treat that field as opaque continuation data: it is not meant to be rendered directly in the UI. When a readable summary is available, that is the user-facing form to show.