Configuration & Auth
Connect your models once, then reuse pools, fallback, and project overrides. Chord separates behavior configuration and credentials:
~/.config/chord/config.yaml: providers, models, extensions, defaults~/.config/chord/auth.yaml: API keys / OAuth credentials.chord/config.yaml: project-level overrides~/.config/chord/agents/and.chord/agents/: agent role definitions
How to use this page
Section titled “How to use this page”You do not need to read this page from top to bottom:
- First setup: start with Quickstart, pick a channel in Choosing models if you have not already, then copy a provider from Model configuration recipes.
- Credentials and OAuth: jump to
auth.yamlor OAuth. - Routing and reliability: use Model pools, Provider timeouts, and Stream retry cap.
- Long sessions: use Context management.
- Exact field names: use the Configuration cheatsheet.
Configuration layers
Section titled “Configuration layers”A practical precedence model is:
- Built-in defaults
- Global config
- Project config
- Agent-level config
This lets you keep personal defaults, project-specific behavior, and per-agent capabilities separate.
Project configuration is read from .chord/config.yaml in the startup directory; Chord does not search parent directories. Set only the fields you want to override and leave the rest inherited. Merge rules:
- omitted project fields stay truly unset instead of silently shadowing global defaults;
- unrecognized keys, wrongly typed values, and out-of-range settings in any config file are logged to
chord.logand treated as not configured while the rest of the file still applies; in a project config an invalid leaf falls back to the inherited global value. Malformed YAML (a syntax error) prevents startup; runchord doctor configfor the full problem list; - settings that load as written but likely do not work as intended, such as a thinking block a Chat Completions gateway never receives, are logged to
chord.logas warnings and listed bychord doctor configwithout counting as problems; - global-only keys such as
paths.*andmaintenance.*(alsomodel_templatesanddiagnostics) are ignored in project config; - most scalar and object values override the global value at the same key;
model_poolsmerge by pool name, with same-name project pools overriding the global definition;mcpmerges by server name, with each same-name project server replacing the entire global server definition rather than inheriting individual connection or permission fields;- append-style extension points keep global entries and add project entries: currently
skills.pathsand per-trigger hook arrays underhooks.*append rather than replace.
If global config.yaml is missing, the first chord run starts a one-time setup wizard that writes config.yaml and, when needed, auth.yaml; see Quickstart. Prefer to write both files yourself? Everything below is the full field reference.
Minimal provider config
Section titled “Minimal provider config”OpenRouter
Section titled “OpenRouter”providers: openrouter: type: chat-completions api_url: https://openrouter.ai/api/v1/chat/completions models: openai/gpt-5.5: limit: context: 400000 input: 272000 output: 128000 modalities: input: [text, image, pdf]BigModel Chat Completions (Coding Plan)
Section titled “BigModel Chat Completions (Coding Plan)”Chord’s api_url is the complete request URL. For the BigModel Coding Plan
OpenAI-compatible endpoint, append /chat/completions to the Coding Plan base.
providers: bigmodel: type: chat-completions api_url: https://open.bigmodel.cn/api/coding/paas/v4/chat/completions models: glm-5.3: &bigmodel-glm-5-3 limit: context: 1000000 output: 128000 reasoning: effort: max compat: request_overrides: rename_body_fields: max_completion_tokens: max_tokens body: thinking: type: enabled clear_thinking: false reasoning_continuity: mode: openai_visible
glm-5.3-flash: <<: *bigmodel-glm-5-3 modalities: input: [text, image, pdf]glm-5.3 is the text-only flagship; glm-5.3-flash takes the same text
parameters and settings and adds image/PDF input, so pick the model ID your
workload needs.
OpenAI Responses
Section titled “OpenAI Responses”For provider/model-specific copy-paste snippets (GPT-5.4/5.5/5.6, Claude, Gemini, GLM, DeepSeek/OpenAI-compatible), see Model configuration recipes.
providers: openai: type: responses api_url: https://api.openai.com/v1/responses models: gpt-5.6-sol: limit: context: 1050000 output: 128000 reasoning: effort: medium summary: auto variants: low: reasoning: effort: low high: reasoning: effort: high xhigh: reasoning: effort: xhigh max: reasoning: effort: max modalities: input: [text, image, pdf]
model_pools: default: - openai/gpt-5.6-sol@xhighPair this provider with an API key in ~/.config/chord/auth.yaml:
openai: - "$OPENAI_API_KEY"gpt-5.6-terraandgpt-5.6-lunaare wired the same way; keep themodelskey and themodel_poolsref on the same model ID.- This snippet targets the official OpenAI API, so it declares the full
1050000window with noinput: Chord then derives the usable input budget ascontextminus the model’s ownoutputcap (1050000 − 128000 = 922000), and reserves the default64000output cap only for models that declare nolimit.output. Above 272K is a pricing threshold here, not an input cap, so do not addinput: 272000. - Codex OAuth uses the same model windows as the API: GPT-5.4 / 5.6 / 6 run
the
1050000 / 922000 / 128000allocation there too (see OpenAI Codex preset below). - Supported API reasoning efforts are
none,low,medium,high,xhigh, andmax; select a configured variant with a ref such asopenai/gpt-5.6-sol@xhigh. - When reasoning is active, Responses defaults
reasoning.summarytoauto; set it tononeto opt out explicitly. Chord does not currently expose GPT-5.6reasoning.mode: pro. preset: codexproviders can also usemaxwhen the selected model/backend supports it. Whether a given effort level is accepted is model/provider-specific.
Limits and reasoning values were checked against the current Codex model catalog and OpenAI’s GPT-5.6 model guidance. Actual relay allocations may be smaller and should follow the gateway’s published model catalog.
Read model limits in this order:
limit.contextis the total window. For most models, input + requested output just needs to fit inside this number.limit.inputis only needed when the provider also lists a separate input cap. Some GPT models work this way; if you omit it, Chord derives the usable input budget aslimit.contextminus the model’s ownlimit.output(only a model declaring no output cap falls back to the globalmax_output_tokensdefault). A declaredlimit.inputis always used as-is.limit.outputis the model’s own output capacity. Chord’s default requested output cap (max_output_tokens) is64000, so real requests usemin(64000, limit.output)before the available-context clamp. Setmax_output_tokensexplicitly to choose a different global cap. If a model’s real output capacity is below64000andlimit.outputis omitted, backends that validate the requestedmax_tokensserver-side will reject those requests. Declarelimit.outputfor such models, or lower the globalmax_output_tokens.
parallel_tool_calls defaults to true for Responses and Chat Completions providers. Set it to false on a provider, model, or variant only when the backend or workflow requires serial tool calls. Provider-level user_agent is also available for gateways that require a specific client identifier.
Provider auth headers are inferred separately from type, but can be overridden with auth_scheme when a compatible endpoint expects a different credential header:
type: messages→ defaultauth_scheme: anthropic-api-key(x-api-key)type: responses→ defaultauth_scheme: bearer(Authorization: Bearer)type: chat-completions→ defaultauth_scheme: bearer
For an Azure OpenAI Responses endpoint, configure a plain type: responses provider with auth_scheme: api-key and store: true, and remove the Codex identity headers with compat.request_overrides.headers set to null (see below).
Supported auth_scheme values are:
anthropic-api-keybearerapi-key
Use an explicit override only when the endpoint’s auth requirements differ from Chord’s transport default. For example, a provider may expose an Anthropic-compatible /messages path but require Authorization: Bearer instead of x-api-key. In that case, keep the same type and set only auth_scheme: bearer.
For Anthropic’s gated 1M context beta, Chord opts in only when the model declares a window of at least 1M tokens (limit.input when set, otherwise limit.context). Models with smaller declared windows do not receive the beta header. The provider may apply different access requirements and pricing above 200K tokens.
store controls whether a Responses backend retains requests and responses server-side. It defaults to false. Enable it only when the backend explicitly requires server-side retention and you accept the data-retention trade-off. Do not enable it for preset: codex; the official Codex OAuth endpoint rejects store: true.
OpenAI Codex preset
Section titled “OpenAI Codex preset”Codex OAuth uses the same model windows as the API examples.
providers: codex: preset: codex type: responses models: gpt-5.5: limit: context: 400000 input: 272000 output: 128000 gpt-5.4: limit: context: 1050000 input: 922000 output: 128000 gpt-5.6-sol: limit: context: 1050000 input: 922000 output: 128000GPT-5.4 / 5.6 Sol / Terra / Luna / GPT-6 Sol / Luna / Astra / GPT-6.1 Sol use 1050000 / 922000 / 128000
(1.05M total window; the 922K input budget derives as context − output,
since these models publish no separate input cap); GPT-5.5 and
GPT-5.2 use 400000 / 272000 / 128000. See Model configuration recipes
for complete examples.
preset: codex can use OpenAI / ChatGPT OAuth credentials from auth.yaml. OAuth entries are mappings:
codex: - refresh: rfr_... access: eyJ... expires: 1774009702606 account_id: acc_... # optional; only workspace/account tokens always have this account_user_id: u_...__acc_... # optional; parsed/backfilled in the background when missing email: user@example.com # optionalaccount_id is not present for every ChatGPT account. Personal Plus/Pro access tokens may carry only user_id and no chatgpt_account_id; Chord still uses those credentials as ordinary OAuth bearer tokens and omits the ChatGPT-Account-ID header. Features that require a workspace/account id, such as Codex usage / rate-limit polling, skip those credentials until an account id is provided or parsed later.
For large account pools, Chord does not synchronously parse every OAuth JWT at startup and does not block provider initialization when one access token lacks account_id. Startup reads only metadata already present in auth.yaml; missing account_user_id, account_id, email, and expires are parsed and backfilled in the background after the provider is available. When manually converting Codex / sub2api / other login exports, keep any available account_id, account_user_id, and email, but they are not startup requirements.
Azure OpenAI Responses
Section titled “Azure OpenAI Responses”Configure an Azure OpenAI Responses endpoint as a plain type: responses provider with auth_scheme: api-key and store: true. Omit the Codex streaming identity headers with compat.request_overrides.headers set to null:
providers: azure: type: responses api_url: https://YOUR-RESOURCE.openai.azure.com/openai/v1/responses auth_scheme: api-key store: true trust_http_400: true retry_after_max_s: 86400 compat: request_overrides: headers: OpenAI-Beta: null originator: null models: gpt-5.5: limit: context: 400000 input: 272000 output: 128000Azure’s v1 Responses endpoint can use /openai/v1/responses directly; add api-version only when you need to pin or opt into a specific version such as preview. Store the Azure API key under the same provider name in auth.yaml:
azure: - $AZURE_OPENAI_API_KEYGoogle Gemini
Section titled “Google Gemini”providers: gemini: api_url: https://generativelanguage.googleapis.com/v1beta/models models: gemini-3.8-flash: limit: context: 1048576 output: 65536 modalities: input: [text, image, pdf]For Gemini, set api_url to the /models base path. Chord detects type: generate-content from the URL path’s /models suffix, so type can be omitted. Do not include the model name or :streamGenerateContent?alt=sse; Chord appends /{model}:streamGenerateContent?alt=sse automatically. The model map key, such as gemini-3.8-flash, is the model ID sent to Gemini.
Gemini thinking options use the same unified thinking object as other providers (no separate gemini_thinking key):
thinking.budget→generationConfig.thinkingConfig.thinkingBudget- Gemini: ✅ used
- Anthropic: ⚠️ only when
thinking.type: enabled(mapped to Anthropic budget mode) - OpenAI: ❌ ignored
thinking.include_thoughts→generationConfig.thinkingConfig.includeThoughts- Gemini: ✅ used
- Anthropic / OpenAI: ❌ ignored
thinking.level→generationConfig.thinkingConfig.thinkingLevel(minimal|low|medium|high, Gemini 3+; not all models supportminimal)- Gemini (3+): ✅ used
- Gemini 2.x / Anthropic / OpenAI: ❌ ignored
Example:
providers: gemini: api_url: https://generativelanguage.googleapis.com/v1beta/models models: gemini-2.5-flash: limit: context: 1048576 output: 65536 modalities: input: [text, image, pdf] thinking: budget: -1 include_thoughts: true gemini-3-pro: limit: context: 1048576 output: 65536 modalities: input: [text, image, pdf] thinking: budget: -1 level: highIf type is omitted, Chord auto-detects it from provider config:
preset: codex→responsesapi_urlpath ending in/responses→responsesapi_urlpath ending in/chat/completions→chat-completionsapi_urlpath ending in/messages→messagesapi_urlpath ending in/models→generate-content
If none of these rules match, set type explicitly.
Appended thinking translation
Section titled “Appended thinking translation”If your model outputs English thinking / reasoning and you want an appended translation (for example, Chinese) in the TUI, you can enable thinking_translation:
model_pools: translation: - openai/gpt-5.4-mini
thinking_translation: target_language: zh-Hans model_pool: translation max_chars: 1000Notes:
- Only thinking / reasoning output is translated, never the assistant final answer. The translation is appended under the corresponding thinking card with a neutral
Translated · <target_language>header, rendered through the same Markdown / code-highlighting pipeline, and never written back into model context. target_languageandmodel_poolare both required; if either is missing, the feature is disabled.model_poolmust point to a top-levelmodel_poolsentry: prefer a separate low-cost translation pool. The pool can contain multipleprovider/model[@variant]refs; translation runs a single fallback round across them in order, moving to the next candidate on failure (including network/5xx/timeout) or when a result is empty, clearly truncated, or in the wrong language.max_chars(default1000) limits the thinking preview sent for translation; only the leadingmax_charsrunes are translated, and text past that prefix will not appear in the translated card. Set a smaller value such as500for lower latency/cost, or a larger one for more complete translations.- A temporary failure only skips that one thinking block; it does not block later thinking translations or the main response. Per-provider transport timeouts (one-minute-class by default) still apply, so a stalled model or key can fail over while the rest of the pool gets a chance to run.
- Translations are persisted in the session directory (
thinking_translations.json) and restored when the session is resumed. A given thinking block is translated at most once: changingthinking_translation.target_languagelater does not re-translate already-stored blocks.
More detailed fields are described in the config reference below.
auth.yaml
Section titled “auth.yaml”Provider keys must match the provider name in config.yaml:
The first-run wizard can create this file for you. It supports either literal API keys or $ENV_VAR placeholders.
anthropic: - "$ANTHROPIC_API_KEY"
openai: - "$OPENAI_API_KEY"You can list multiple keys for rotation or backup.
For preset: codex OAuth providers, Chord keeps frequently changing runtime status (quota snapshots, reset times, last warm-up timestamps, shared OAuth status cache) in auth.state.json, not in auth.yaml.
That split is intentional:
auth.yamlremains the user-edited source of truth for credentials and stable OAuth fields such asrefresh,access,expires,account_id, andemail; empty OAuth fields are omitted when Chord rewrites the file, and OAuthstatusdoes not belong inauth.yaml;auth.state.jsonis machine-managed shared runtime state. Normal entries are keyed directly byaccount_user_idbelow each provider so quota / reset updates and account states such asexpired,deactivated, andinvalidateddo not constantly rewriteauth.yamlwhile the user may also be editing it. Refresh-only credentials whose account is not known yet can temporarily use arefresh_sha256:<digest>state entry until the first successful refresh backfillsaccount_user_id. State entries without a matchingauth.yamlOAuth credential, and unrecognized legacy state-key formats, are removed bychord auth state clean.
For OAuth credentials with access, the access token must carry parseable account and user/account-user claims. If auth.yaml already has account_id, the token’s account ID must match; otherwise the access token is rejected as a mismatched credential. Chord can also keep a refresh-only OAuth entry (refresh without access) and refresh it on first use; after a successful refresh, Chord extracts account_id and switches runtime state to the account_user_id key.
If refresh fails unrecoverably before the account is known, Chord records the invalid state under refresh_sha256:<digest> so chord auth state clean can remove the unusable credential later. An OAuth entry with neither access nor refresh is unusable.
expires is the access-token expiry timestamp in Unix milliseconds. When access contains a JWT exp claim, Chord uses that value as the most accurate expiry metadata and can cache the resulting expiry in auth.state.json without storing the access token there. A missing or locally expired expires value does not by itself mark an OAuth slot expired or unhealthy. Chord still tries the existing access token first, and only after an authentication failure will it refresh the credential or mark it expired if recovery is impossible.
Typical auth.state.json content looks like:
{ "openai": { "user-1__acc-1": { "account_user_id": "user-1__acc-1", "account_id": "acc-1", "email": "user@example.com", "expires": 1774009702606, "status": "expired", "updated_at": 1774009702606, "last_warmup_at": 1774009702606, "codex_primary_used_pct": 12.5, "codex_primary_window_minutes": 60, "codex_primary_reset_at": 1774013302000, "codex_secondary_used_pct": 40, "codex_secondary_window_minutes": 10080, "codex_secondary_reset_at": 1774600000000 } }}The status field is authoritative only in auth.state.json. Chord writes expired when an access token can no longer be used and the credential cannot be refreshed (including missing, invalid, expired, or reused refresh tokens), deactivated when the service reports a disabled/banned account, and invalidated when the account must be re-authenticated. Any non-empty status makes that OAuth slot unselectable until it is cleaned up or replaced.
These cached Codex quota/reset fields are restart-stable scheduling and display hints, not hard blocks by themselves:
- they help startup / first-pick ordering choose accounts that are more likely to still have quota;
- they let key switches immediately show the last cached snapshot before a fresh warm-up completes;
- they do not by themselves make the account absolutely unselectable;
- real hard blocking still comes from confirmed request failures and runtime cooldown state.
Environment variables in auth.yaml
Section titled “Environment variables in auth.yaml”Provider credentials in auth.yaml support environment-variable expansion for scalar API-key values:
anthropic: - "$ANTHROPIC_API_KEY"
openai: - "${OPENAI_API_KEY}"Expansion is applied when the scalar starts with $. Unset variables expand to an empty string and are filtered out, unless the YAML value is a literal empty string. This expansion applies to auth.yaml credentials, not generally to every field in config.yaml.
If you intentionally need an empty API key, write a literal empty string:
local-provider: - ""Do not rely on an unset environment variable for this case. An unset $ENV_VAR is treated as a missing credential and is filtered out.
Provider key selection
Section titled “Provider key selection”When a provider has multiple API keys / OAuth accounts, Chord uses two settings: key_rotation controls when Chord reselects a key, and key_order controls how Chord chooses among selectable keys.
key_rotation: on_failure(default): keep using the current key until it fails, cools down, or becomes unusable.key_rotation: per_request: reselect a key before every request; useful for load balancing across independent keys.key_order: sequential(default for non-Codex providers): choose in stable key order, generally preferring the least-recently-used selectable key.key_order: random: choose randomly among selectable keys.key_order: smart: Codex providers only. Prefer healthy OAuth accounts with better quota headroom and reset timing.
key_rotation only rotates credentials / API keys. It does not rotate models; model selection still follows the model pool sticky cursor and fallback logic.
Loop mode still follows the configured key_rotation / key_order. For Codex long-running loops, keep the default key_rotation: on_failure if you prefer stable transport/cache continuity; explicitly use per_request only when you want to distribute quota across multiple accounts.
Only providers with preset: codex are treated as OAuth providers.
For Codex providers, prefer configuring only preset: codex plus model settings. Do not manually override preset-managed fields such as api_url, token_url, client_id, type, store, responses_websocket, or supported_service_tiers unless you are deliberately testing transport internals. The preset selects the official OAuth transport, Responses endpoint, WebSocket/cache defaults, quota polling, smart key ordering, and service-tier capability.
It does not define a separate HTTP request body or force a Codex User-Agent: non-Codex type: responses providers use the same Responses wire shape described above, and all providers default to User-Agent: chord/<version>. Use supported_service_tiers when you need an explicit tier matrix.
Codex OAuth account selection is controlled by key_rotation / key_order in Provider key selection. Codex defaults to key_order: smart, which considers quota snapshots, soft cooldown, and reset timing when choosing an account.
smart ranks selectable Codex OAuth accounts by preferring:
- accounts whose cached snapshot has no tracked window at
100%used; a100%window is tried last but is not a hard block by itself; - accounts with remaining quota in the shorter primary window (for example the 5h window), choosing the nearer primary reset first so soon-expiring quota is used before it is wasted;
- then accounts with remaining quota in the longer secondary window (for example the 1w window), again preferring the nearer secondary reset;
- then higher remaining headroom when the comparable windows reset at the same time;
- still falling back to unknown / stale candidates when no better option exists.
When a Codex client becomes active, Chord may also background-probe additional OAuth slots to refresh cached headroom snapshots. That warm-up is best-effort, low-concurrency, cancels when the active client is replaced, and only refreshes cached quota state; authentication failures from usage probes do not mark OAuth credentials unusable.
Warm-up priority is also state-aware:
- OAuth slots that have never been warmed up in shared state are probed first;
- older cached entries are refreshed before recently refreshed ones;
- after warm-up or polling returns a newer snapshot, Chord writes it to
auth.state.jsonand other processes adopt it lazily when they next read key-selection or rate-limit state.
# auto-select a configured codex providerchord auth
# explicitly choose a providerchord auth codex
# headless / SSH environmentschord auth codex --device-codeModel pools (selecting provider/model)
Section titled “Model pools (selecting provider/model)”Chord selects the active model via named model pools. Each pool entry should be a full provider/model[@variant] reference so the provider endpoint, auth, protocol, and variant tuning are unambiguous.
Pool definitions live in config.yaml (global or project-level). Agent configs
may reference pool names to restrict access; they cannot define inline pools.
Define model pools in config.yaml
Section titled “Define model pools in config.yaml”# ~/.config/chord/config.yaml or .chord/config.yamlmodel_pools: thinking: - anthropic/claude-opus-5 - openai/gpt-5.5 non-thinking: - anthropic/claude-sonnet-4Project-level .chord/config.yaml model_pools are merged into the global config
(same-name pools override).
Reference pools from agents (optional)
Section titled “Reference pools from agents (optional)”Agents do not need to set model_pools. If omitted, the agent can use every
pool defined in merged config.yaml model_pools, sorted by pool name. Add
model_pools: [...] only when you want to restrict that agent to a subset or
customize its fallback order.
# ~/.config/chord/agents/builder.yaml or .chord/agents/builder.yamlname: buildermode: mainmodel_pools: [thinking, non-thinking]name: reviewermode: subagentmodel_pools: [thinking]When no pool is explicitly selected, Chord falls back to the agent’s first
allowed pool: the first entry in model_pools: [...] when configured, otherwise
the alphabetically first top-level pool.
At runtime, use /models to switch the pool for the current view (per project,
persisted across restarts). In the main view this means the current main role; in a
SubAgent view it means that SubAgent’s agent pool selection.
Switching pools updates
the full fallback chain for subsequent LLM calls, even if the currently selected
provider/model exists in both pools (in-flight requests keep using their starting
snapshot). You can also set a named
agent directly with /models --agent <name> <pool>. For SubAgents, the default behavior
is to use the first allowed pool; switching back to that pool restores the default
behavior.
User messages submitted while the main agent is busy remain queued until a safe request boundary. If the current provider/model attempt fails and Chord is about to send a request to the next fallback model, messages queued by that point are committed to the conversation and included in that fallback request. Chord never rewrites a provider request that is already in flight; messages arriving after the fallback request starts wait for the next request boundary.
Reusing protocol templates with YAML anchors
Section titled “Reusing protocol templates with YAML anchors”Chord has no model_templates schema field. You can still use YAML anchors and
merge keys under that top-level container; Chord ignores the container itself
and reads the expanded model entries under providers.
Merge keys (<<:) copy the referenced mapping into the current entry at the
key level, and the current entry wins on conflict:
- Scalar fields (
reasoning.effort, acompactionfraction) simply replace the inherited value. - Nested objects are replaced as a whole, not merged field by field:
overriding with
limit: {context: 1050000}on a template that already declaredlimit: {context: 1000000, output: 128000}silently drops the inheritedoutput. Write the complete block for anything you override (limit,cost,compaction,variantsentries,modalities, …). - The same rule holds along a chain (
gpt-5.6-luna: &gpt-5-6-luna {<<: *gpt-5-6-base}): the deepest entry wins per whole key. - You cannot unset a field an ancestor template declares: override it with
a concrete value, or stop referencing that template.
compaction.threshold: 0andcompaction.reminder: -1are the documented exceptions that disable those two behaviors explicitly.
This page covers protocol and field semantics. For current model limits, pricing, and complete GPT / Claude / Gemini / GLM / DeepSeek snippets, see Model configuration recipes.
model_templates: chat-thinking: &chat-thinking limit: context: 200000 output: 64000 reasoning: effort: high compat: # DeepSeek and GLM Chat APIs document max_tokens, not the OpenAI # reasoning-model default max_completion_tokens. request_overrides: rename_body_fields: max_completion_tokens: max_tokens body: # Gateways for the model families listed in "Thinking behind a Chat # Completions gateway" (model-configs.md) get this object written # automatically; the override covers other backends — and any extra # field of the same object, such as GLM's clear_thinking. thinking: type: enabled # Replays native reasoning_content and accepts portable visible # reasoning from other wire families without injecting request fields. reasoning_continuity: mode: openai_visible
responses-thinking: &responses-thinking limit: context: 200000 output: 64000 reasoning: effort: high summary: auto
messages-thinking: &messages-thinking limit: context: 200000 output: 64000 thinking: type: adaptive effort: high compat: # Compatible endpoints should not receive Anthropic beta headers unless # their own documentation opts into them. request_overrides: headers: anthropic-beta: null
providers: chat: type: chat-completions api_url: https://example.com/v1/chat/completions models: chat-model: *chat-thinking
responses: type: responses api_url: https://example.com/v1/responses models: responses-model: *responses-thinking
messages: type: messages api_url: https://example.com/v1/messages models: messages-model: *messages-thinkingModel field semantics:
limit.context: total request window when the provider publishes one. If it is omitted and bothlimit.inputandlimit.outputare positive, Chord derives it asinput + output; an explicitcontextalways takes priority.limit.input: independent input cap when published. A declared value is authoritative and used as-is, even when it is not additive withlimit.outputinside the window. If omitted, Chord derives the prompt budget aslimit.contextminus the model’slimit.output(so the 1.05M GPT family gets 1050000 − 128000 = 922000); only a model declaring no output cap falls back to reserving the effective default output cap (max_output_tokens, default64000).limit.output: model output capacity. Runtime requests are also capped by the globalmax_output_tokenssetting and remaining total-context space.reasoning.effort: reasoning depth. Chord keeps no local whitelist: whatever level the provider supports reaches the upstream unchanged, and the Responses wire additionally normalizes whitespace and casing before sending.- Chat Completions sends top-level
reasoning_effort. - Responses sends
reasoning.effortand optionalreasoning.summary.
- Chat Completions sends top-level
reasoning.effort_map: maps the canonical effort value to the wire value the provider actually accepts, for example{high: max}when a gateway exposesmaxfor Chord’shigh. The mapping applies to the final resolved effort, so variant-level maps replace the model-level map for that variant.reasoning.summary: Responses reasoning summary request. Supported Chord values areauto,concise,detailed, andnone. When reasoning is active, omission defaults toautoso cross-provider replay retains portable summary text; usenoneto opt out explicitly.thinking: Messages-compatible extended thinking.type: adaptivecombines withthinking.effort, which Chord sends asoutput_config.effort.text.verbosity: optional OpenAI-compatible visible-text verbosity hint.variants: named model parameter overrides selected with refs such asprovider/model@high.cost: optional USD-per-million-token estimates. It can include input/output, cache prices, service-tier multipliers, and long-context input tiers.modalities.input: supported input kinds:text,image, andpdf.supported_service_tiers: accepted non-standard tiers such asfastorslow; price multipliers are configured separately undercost.
Compatibility fields:
-
compat.request_overrides.body: recursively merges arbitrary JSON into the final protocol request. Anullvalue deletes that field. -
compat.request_overrides.rename_body_fields: renames a final request field while preserving Chord’s dynamically computed value. Use this for differences such asmax_completion_tokens: max_tokens. -
compat.request_overrides.headers: sets arbitrary request headers. Anullvalue removes a Chord default header, for exampleanthropic-beta: nullon a compatible Messages endpoint. -
compat.reasoning_continuity.mode:-
none: no provider-specific visible reasoning replay. -
openai_visible: replays unchanged assistantreasoning_contentduring Chat Completions tool loops and accepts portable visible reasoning from other wire families asreasoning_content. It does not inject request fields; configure those withrequest_overrides.body. On the first attempt Chord still optimistically replays chat-native reasoning to any Chat Completions target, even across providers, so documented in-provider upgrades such as Kimi K2.6/K2.7 to K3 and same-model provider fallback both keep continuity.If a target rejects that request, Chord degrades the target for the rest of the session while keeping completed tool calls and paired results structured until strict compatibility requires text facts. On
openai_visibleResponses targets whose thinking mode requires replayed function-call turns to carry reasoning, Chord replays the native plaintextreasoning_textwhen available.For turns that lost their native reasoning (for example after a cross-provider model switch), a backend rejection escalates the replay to plain-text historical tool records, so the continuation no longer needs the missing reasoning. For third-party OpenAI-compatible gateways, Chord keeps the configured endpoint and does not redirect to DeepSeek’s official
/betaendpoint.A reasoning-only output truncation may receive one bounded request-only reasoning replay on that same endpoint; if the gateway rejects it, Chord falls back to the ordinary recovery prompt.
-
anthropic_unsigned: opt-in for Messages-compatible models such as DeepSeek/GLM endpoints that return visiblethinkingblocks without Claude signatures. Unsigned thinking is replayed natively only to the same provider/model on the first attempt. For compatible targets, portable visible reasoning from other wire families is converted into unsignedthinkingblocks instead of being injected into assistant text. -
Responses, signed Claude Messages, and Gemini otherwise use their protocol-native continuity mechanisms automatically. Chord captures opaque encrypted/signature state and replays it only where the target wire and provenance permit. Across incompatible protocols, portable visible reasoning is converted only when the target has a structured carrier (
openai_visibleoranthropic_unsigned); otherwise it is dropped. Opaque state is never fabricated, and reasoning is never injected into ordinary assistant content.Completed tool facts are converted to the target protocol’s structured representation whenever possible and textified only if a target rejects that shape. The achieved degradation level is remembered per target.
-
-
compat.reasoning_continuity.reasoning_replay: selects how much reasoning Chord replays from completed turns (everything before the last user message). DeepSeek Chat/Messages always preserve the complete reasoning history; for other targets the defaultcurrent_turnstrips completed-turn reasoning,allreplays it unchanged, andnonealso strips the current turn.Completed turns are stripped by default to keep the request small. Anthropic filters prior-turn thinking blocks server-side and bills only the blocks the model actually sees, so omitting them costs nothing there; targets that replay history verbatim charge for every retained token, which is why the contracts listed below opt into
all. The policy covers every reasoning payload — plaintextreasoning_content, unsignedthinkingblocks, and provider-bound native thinking (signed Claude blocks, Responses reasoning items, Gemini thought signatures) — while the tool trajectory, including call/result pairing, always survives the strip.Set
allwhen the target’s contract requires the complete assistant history (Kimi K3 andkeep: allmodels, Qwenpreserve_thinking, GLMclear_thinking: false, Xiaomi MiMo; DeepSeek Chat/Messages already keep it without this setting); historical reasoning is then replayed unchanged and billed on every request. Setnoneonly for an endpoint you verified accepts a request without the current turn’s reasoning. Current-turn reasoning — everything after the last user message, including its tool loop — is otherwise replayed unchanged.Anthropic additionally binds each thinking block to the conversation prefix that produced it: when a history rewrite invalidates that binding and the API rejects the replay with an invalid-signature error, Chord retries once with the thinking blocks dropped and keeps the turn’s text and completed tool facts.
Request-scoped turn overlays (per-turn
<system-reminder>hints) are not counted as user turns, so an overlay appended at the tail cannot shift the completed-turn boundary past the current turn and strip the reasoning the backend consumes in this turn’s tool chain. -
compat.forced_tool_choice.suppress_in_thinking: downgrades loop-forcedtool_choice: requiredto the backend default while reasoning/thinking is active. Enable it only for OpenAI-compatible endpoints that reject forced tool choice in thinking mode; ordinary tool availability and automatic tool choice still work. -
compat.forced_tool_choice.auto_only: downgrades any non-autotool_choiceto the backend default unconditionally. Enable it for backends that only supporttool_choice: "auto"and rejectrequired,none, or named choices with a 400; automatic tool choice still works. Chat Completions, Messages, and Gemini omit the field; Responses still sends an explicittool_choice: "auto"unlesscompat.responses.send_tool_choiceis false. -
compat.thinking_toolcall: enables a provider-specific parser for gateways that encode tool calls inside visible reasoning text. Leave disabled unless the gateway requires that format.
Provider-level compat values are defaults. A model-level compat block can
override them for one model.
Request overrides apply to HTTP transports after Chord has constructed the protocol request. Configuring them on a Codex Responses provider disables its WebSocket transport for that request so the final JSON patch can be honored.
Project-level config
Section titled “Project-level config”If a project needs local defaults, create this file at the project root:
.chord/config.yamlCommon uses include:
- Project-specific permission rules
- Project-specific LSP / MCP / Hooks / Skills settings
Provider request compression
Section titled “Provider request compression”Provider-level compress selects the encoding for compressed upstream request
bodies: gzip or zstd (zstd is the codec the Codex client uses for
codex-backend request bodies). It is different from context management
(compaction / reduction): it only changes HTTP request transfer encoding and
does not summarize or remove conversation history.
providers: openai: compress: gzip codex: preset: codex compress: zstd # codex-backend accepts zstd request bodiesChord compresses the request body only if compression reduces the payload
size; otherwise it sends the request uncompressed. Compression failures are
logged and the request is sent uncompressed too. The response direction is
unaffected: Chord still advertises only gzip responses and decodes them
itself.
Provider/model requests identify the client with User-Agent: chord/<version> by default. Set provider-level user_agent only when a provider or gateway requires a specific value:
providers: gateway: user_agent: RequiredGatewayClient/1.0This setting also applies to Responses HTTP requests, Codex OAuth requests, Codex usage polling, and Responses WebSocket handshakes for that provider. WebFetch uses its own web_fetch.user_agent.
To send a Codex-style User-Agent, set the captured Codex value explicitly and keep it updated when the Codex client version or terminal changes:
providers: codex: preset: codex user_agent: "codex-tui/0.139.0 (Mac OS 15.3.2; arm64) ghostty/1.3.1 (codex-tui; 0.139.0)"Provider retry delays
Section titled “Provider retry delays”Provider-level retry settings control the generated delay between complete
retry rounds. A round starts at the current sticky model-pool cursor and walks
the remaining pool entries and their eligible keys without sleeping between
fallback targets. The cursor provider’s settings own the delay before the next
round; if a fallback later succeeds and becomes the sticky cursor, its settings
apply to subsequent requests. The same settings also control ordinary HTTP 429
key cooldown when the response carries no Retry-After hint; a valid hint
(bounded by retry_after_max_s) always outranks them.
providers: gateway: retry_backoff: exponential retry_delay_ms: 500retry_backoff: exponentialis the default. It starts atretry_delay_ms(default1000) and doubles each round, capped at 60 seconds.retry_backoff: fixedusesretry_delay_msfor every round.retry_backoff: nonedisables generated round backoff. An ordinary 429 without aRetry-Afterhint marks the failed key as recovering without a timed cooldown, so healthy keys and fallback targets still take priority. Rounds are unlimited by default, so withnonea fast-failing upstream is retried back-to-back without delay.retry_delay_msaccepts0through60000;0/ omitted means 1000ms. Invalid modes, negative values, and values above the cap are configuration errors.
A retry round that pauses before its next attempt shows that wait in the status
bar as a countdown to the attempt (↺ round 12 · retry in 45s). A round with no
delay re-probes immediately and keeps showing how long the retry has been
running, since it has no wait to count down.
For an ordinary 429, the key cooldown follows a single priority order: a confirmed quota reset window wins, then a valid Retry-After (bounded by retry_after_max_s) applies verbatim, and only a hint-less 429 falls to the retry pacing above: the configured exponential/fixed/none mode, or the one-second exponential default when neither field is set. Invalid or deactivated credentials, and cooldowns already established by other hard states, are never shortened or cleared.
This 429 pacing applies before and after visible streaming output alike: a 429 that interrupts a visible stream cools the key down and rotates to the next one.
When every key of every pool entry is cooling down, Chord waits instead of
sending a request, and how long it sleeps depends on the pool. With a single
model configured there is nothing else to try, so it waits for the earliest key
recovery instant — a confirmed quota reset instant the provider defines, or a
Retry-After hint already capped by retry_after_max_s. Nothing re-probes the
pool during that stretch, so a credential added mid-wait is picked up once the
wait ends. With fallback models configured it re-checks the pool at least once a
minute, because a sibling model, a newly added credential, or a refreshed
rate-limit snapshot can free up a request long before the longest cooldown ends.
Either way the status bar counts down to the point a request can actually go
out, not to the next internal re-check. The shortest cooldown in the pool always
decides: a model that is ready again is never held back by a longer cooldown on
another one, and a pool that is re-checked every minute picks a recovered key up
as soon as it is ready.
When a key goes into cooldown, the API failure behind it is recorded in the
error panel (Ctrl+E) by the wait that shows the cooling or by the next
foreground request (an agent turn or a context compaction), whichever comes
first. This includes failures from background work such as memory extraction
or thinking translation. A credential the provider permanently invalidated (an
expired refresh token, an invalidated or deactivated account) is announced by
the next foreground request. Each failure is recorded once: a failure already
reported as a retry error by the attempt that caused it is not repeated.
Codex OAuth follows the same rules: every Codex 429 is an ordinary 429. A
retry hint (Retry-After or WebSocket resets_in_seconds) is honored ahead
of explicit settings, and a usage-limit 429 carrying neither a hint nor an
exhausted quota snapshot uses the ordinary defaults above, not the one-minute
cooldown of the codex preset, which still covers non-429 usage-limit errors.
When a Codex rate-limit snapshot shows an exhausted window with a future
reset, Chord treats that as confirmed quota exhaustion and keeps the provider
reset authoritative.
Project-level .chord/config.yaml can override these fields for one provider.
Provider timeouts
Section titled “Provider timeouts”Provider-level timeout settings are optional and use seconds. Unset or 0 keeps the built-in defaults.
providers: codex: response_header_timeout: 180 stream_idle_timeout: 90 stream_total_timeout: 1800 websocket_handshake_timeout: 45response_header_timeout: timeout from starting a streaming HTTP request until response headers arrive, including connection setup and request-body upload. It stops once headers arrive and does not cap the total duration of a healthy stream; usestream_idle_timeoutto bound gaps between streamed chunks.0keeps the built-in default.stream_idle_timeout: maximum idle time between streamed model data. When set, it overrides both the normal SSE idle timeout and slow-phase idle timeout for that provider, and it also applies to Codex Responses WebSocket reads.stream_total_timeout: wall-clock cap in seconds for one stream, counted from when the response body starts. When set, it also applies to Codex Responses WebSocket reads.0/ omitted keeps the default of no cap: a stream that keeps producing data is slow, not broken, and is bounded bystream_idle_timeoutalone. Set it to bound the one shape the idle timeout cannot catch: a stream that drips data often enough to reset the idle timer but never finishes. The read fails with a timeout error so the normal key/model retry path handles it.websocket_handshake_timeout: Responses WebSocket handshake timeout for providers using that transport, mainlypreset: codexwithresponses_websocketenabled.
These settings are provider-scoped, so project-level .chord/config.yaml can override them for one provider without changing other providers. They do not change fixed low-level connection defaults such as dial or TLS handshake timeouts.
Output token cap
Section titled “Output token cap”Use max_output_tokens to set a global cap on requested output tokens. It defaults to 64000. The effective request limit is still clamped by each model’s limit.output and available total context (limit.context when known), so runtime uses the smallest applicable value across all providers.
Responses providers keep the stable Responses wire shape and do not send a
max_output_tokens field on the HTTP or WebSocket request by default. For a
compatible non-Codex gateway that needs an explicit server-side output cap, set
compat.responses.send_max_output_tokens: true; the other Responses field
toggles are available under the same provider-level object. The global value
still affects Chord-side budgeting and compatibility checks when it is omitted
from the wire request.
limit.input is separate: use it only for models whose providers publish an extra input cap beyond the total context window. Lowering max_output_tokens can reduce cost and long-response failure risk, but it does not increase a provider’s input allowance or replace limit.input.
max_output_tokens: 64000Stream retry cap
Section titled “Stream retry cap”Use stream_retry_rounds to put a hard ceiling on public LLM retry rounds.
Each round can still walk the current model pool and each provider key in the
normal order; this setting limits how many full rounds CompleteStream will
make before giving up.
A “round” here means the whole public retry pass, not a single provider/model
attempt. For example, stream_retry_rounds: 2 allows at most two full passes
through the active routing chain. Once the cap is reached, Chord stops even for
retry classes that would normally wait and continue, such as all-keys-cooling,
concurrent-request 429 responses, or retryable HTTP 400 responses from a
non-official compatible gateway.
Provider HTTP 400 handling is intentionally conservative:
-
Official APIs treat 400 as a terminal invalid-request error.
-
Non-official compatible gateways may return 400 for transient gateway states such as concurrency limits or upstream capacity. Those non request-shaped 400s can cool the current key, rotate to the next key, and continue after all keys are cooling.
-
Request/parameter/model-incompatible 400s still stop instead of retrying forever, for example a structured
code/typesuch asinvalid_request_error,invalid_request, ormissing_required_parameter, or message-only inputs likemissing required parameter,Store must be set to false, orStream must be set to true. -
0keeps the default behavior: retry until success, cancellation, or a terminal failure. -
Positive values stop after that many rounds, even for cooling / concurrent-request retry classes.
-
This is mainly useful for automation or headless environments that prefer bounded latency over maximum persistence.
stream_retry_rounds: 3Local TUI options
Section titled “Local TUI options”These options affect the local TUI. They can be set in the global config and
can also be overridden by project-level .chord/config.yaml when appropriate.
desktop_notification: truedesktop_notification_foreground: trueime_switch_target: com.apple.keylayout.ABCprevent_sleep: true-
desktop_notification: enables terminal notifications in local TUI mode, regardless of whether the terminal is focused. Each notification pairs the terminal notification escape sequence (auto-selected by terminal, OSC 9 or OSC 777) with a terminal bell (BEL), so it can be heard even where the terminal hides notification banners while focused.Chord notifies when the agent actually ran and then stopped (a completed, cancelled, or loop-finished turn, or all SubAgents finishing) and for permission confirmations and questions, Handoff, loop decisions, and notify-protocol corrections waiting for input; user-initiated navigation that settles into idle (session / model-pool / MCP switches, idle slash commands) stays silent. Whether the bell is audible depends on terminal setup; see Platforms.
-
desktop_notification_foreground: controls whether notifications (both the escape sequence and the bell) are sent while the TUI is focused. Defaults totrue; set it tofalseto notify only when the terminal is unfocused. -
ime_switch_target: usesim-select(im-select.exeon Windows) to switch to the specified input method when entering Normal mode, and restore the previous input method when returning to Insert mode. This is useful when you want command keys to use an English keyboard layout. -
prevent_sleep: prevents macOS idle sleep while any agent is active. It is only effective in local TUI mode.
WebFetch
Section titled “WebFetch”web_fetch uses a built-in browser-like User-Agent by default. You can override it in config when a site needs a different header:
web_fetch: user_agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/136.0.0.0 Safari/537.36This setting works in both global config and project-level .chord/config.yaml; project config overrides the global value.
You can also configure a proxy for WebFetch requests:
web_fetch: proxy: socks5://127.0.0.1:1080 # http, https, socks5 supportedproxy: nil(default): inherits the globalproxysettingproxy: ""(empty string): explicitly disables proxy (“direct” mode)proxy: "http://...","https://...","socks5://...": uses specified proxy
web_fetch intentionally remains a lightweight static HTTP reader. It does not run a local browser; JS-heavy pages may be marked as Content-Quality: suspect-shell when the returned HTML looks like an application shell rather than readable content.
WebSearch
Section titled “WebSearch”web_search searches through the provider’s hosted search tool and returns a summary with numbered sources. It is off by default: enable it for an Anthropic Messages or OpenAI Responses provider (or one model) whose endpoint supports hosted search.
providers: anthropic: type: messages compat: hosted_tools: [web_search]On an OpenAI Responses provider, enable it the same way on the responses wire type:
providers: openai: type: responses compat: hosted_tools: [web_search]compat.hosted_tools lists the hosted tools a provider’s models may serve; an omitted model field inherits the provider list, an explicit hosted_tools: [] disables all hosted tools for that model, and a non-empty list replaces the provider list. web_search is a built-in entry, declared as web_search_20250305 on Anthropic Messages and web_search on OpenAI Responses.
The tool joins the model’s tool list while its routing source contains an enabled target that can carry the declaration. Without model_pool, that source is the calling agent’s active model pool; with a named model_pool, it is the configured pool for the hosted request. Otherwise Chord withholds the tool. Each call sends a separate request carrying only the query and declares the hosted search tool there, so the main conversation request never declares it and its history stays free of provider-specific blocks. Chord returns the native results as an ordinary tool result; allowed_domains and blocked_domains travel as request parameters rather than query text.
OpenAI’s search restrictions depend on the model and the sub-request’s reasoning settings: gpt-5 with reasoning.effort: minimal does not support web search, while gpt-5.4 with reasoning.effort: none may produce lower-quality results. See OpenAI’s web search guide for per-model support.
Chord’s Responses search sub-request omits the model’s configured reasoning.effort, so it uses the server’s default reasoning settings. request_overrides can inject a reasoning parameter into that sub-request; the restrictions above apply to the settings actually sent.
The sub-request bills as tokens on the model that serves it. Providers may charge a per-search fee on top; Chord’s cost accounting counts tokens only.
For complete GPT and Claude provider recipes, see Model configuration recipes.
Hosted tools
Section titled “Hosted tools”The top-level hosted_tools section defines provider-side (hosted) tools. Each entry becomes a local tool whose calls run one sub-request declaring the hosted tool, so a tool the provider executes server-side — web search, code execution, file search — needs configuration instead of code. A local tool appears while some provider or model enables it through compat.hosted_tools and the target’s wire type has a declaration in the entry; otherwise Chord withholds it.
Execution happens on the endpoint that serves each sub-request, so availability follows the channel rather than the model alone: a relay can serve the same models without carrying their hosted tools, and some hosted tools may be available only through the provider’s official channels. A declaration the endpoint rejects or silently drops fails the call as described for compat.hosted_tools.
hosted_tools: code_execution: description: Run Python code in the provider's sandbox and report its output. parameters: type: object properties: code: type: string description: Python source to run. required: [code] prompt: "Run this code and report the result:\n{code}" read_only: true timeout_s: 300 model_pool: tools declarations: messages: tool: {type: code_execution_20250825, name: code_execution} force: {type: tool, name: code_execution}| Field | Default | Description |
|---|---|---|
description |
(empty) | Local tool description shown to the model. |
parameters |
object schema without arguments | JSON Schema of the local tool’s arguments. |
prompt |
arguments as JSON | Sub-request instruction; {arg} is replaced with the matching argument value. |
read_only |
false |
Marks the tool read-only for scheduling; the role permission rules still decide access. |
concurrency_safe |
false |
Allows the call to run alongside other concurrent-safe read-only tools in one batch. |
retry_safe |
false |
Allow retries when execution outcome is unknown; enable only when repeated execution is safe. Built-in web_search defaults to true. |
image_paths |
(empty) | Dot-separated object paths to base64 images in each call result, such as [output.image]. Images become tool result attachments; array traversal is unsupported. |
timeout_s |
120 |
Whole-call budget in seconds, covering every sub-request attempt. |
declarations.<type> |
Wire declaration for one provider type: messages or responses. |
|
declarations.<type>.tool |
(required) | Wire JSON with a non-empty type, inserted into that family’s tools array. Other fields follow the provider’s schema. |
declarations.<type>.force |
(omitted) | Raw tool_choice value that forces the call. Omitted, the tool is still declared and the prompt asks for it, but a model that answers without calling it fails the call. |
declarations.<type>.include |
(omitted) | include selectors, for example [web_search_call.action.sources] on Responses. |
declarations.<type>.headers |
(omitted) | Declaration-specific HTTP headers, applied after provider headers and request overrides. Beta headers can be replaced; authentication, transport, and session headers are protected. |
model_pool |
(empty) | Named model_pools entry that serves this tool’s sub-requests instead of the caller’s own pool. Empty follows the caller: the main agent uses the main pool and a subagent uses its own. The pool must exist at startup and the calling agent must include it in its own model_pools; compat.hosted_tools still decides which entries in the pool can serve the tool. To pin a single model, create a pool containing just it — pool entries accept provider/model@variant. |
Inside a declaration, {"$arg": "<name>"} is replaced with the local argument of that name; a missing or empty argument drops the key, and an object that loses all its keys is dropped too. The built-in web_search entry uses this to pass allowed_domains and blocked_domains as declaration parameters.
An entry that only exists in your configuration starts with conservative traits: not read-only, not concurrency-safe, and no automatic replay of operations with unknown outcomes. Built-in entries (currently web_search) are merged field by field, so a declaration can be retargeted to another tool version without restating the local tool surface.
Names must not collide with registered tools, reserved built-in names, or the mcp_ prefix. Startup validates the merged catalog: each entry needs an object parameter schema, a non-negative timeout, and at least one messages or responses declaration with a non-empty tool type. Headers must have valid names and values, and cannot replace authentication, transport, or session headers. Provider-specific declaration fields remain subject to the endpoint’s validation. The built-in web_search entry supports field-level overrides.
The first sub-request walks capable targets of the tool’s routing source. Unset model_pool follows the calling agent: the main agent walks its active model-pool cursor and a subagent walks its own pool. Each hosted tool remembers its successful target per routing source — shared across agents for a named model_pool, isolated per caller otherwise — and later calls start there while the pool contents stay the same; a pool rebuild or a caller pool switch starts fresh, and targets removed from the pool or no longer capable are not reused. If no target in the routing source can carry the tool, the call fails explicitly and never falls back to the main conversation’s model. For Messages pause_turn, Chord retains the complete ordered native content and sandbox container and continues on the same target, with at most four continuations under the original timeout_s budget. Token usage for all attempts belongs to the calling Agent and turn. An explicit tool_choice rejection gets one retry on the same target without forcing the call.
A named model_pool that is missing or empty prevents startup. Chord skips model entries that cannot be loaded and writes the reason to the logs. If none of the routing pool’s usable models supports the tool’s protocol and enables it through compat.hosted_tools, or the calling agent is not authorized to use the named pool, Chord hides the tool and logs the reason once for that routing source. Check the pool definition, the agent’s model_pools, and the targets’ compat.hosted_tools when a tool is missing.
With the default retry_safe: false, a possibly executed operation whose outcome is unknown (connection interruption, stream error, or exhausted continuation limit) stops automatic key/model replay. A clear rejection before execution may still try another target. Set retry_safe: true only when repeating the operation is safe.
Hosted requests share orchestration.max_active_llm_requests, provider_max_active_requests, and model_max_active_requests with other Agent requests. For a provider that allows only one concurrent request, set its entry under provider_max_active_requests to 1. Limits are local to this Chord process; separate provider entries or processes sharing one account do not share a quota gate.
For retry_safe: true tools, including the built-in web_search, transient rate limits, upstream unavailability, and transport failures get at most three rounds per target, with configured key rotation, provider backoff, and Retry-After pacing. Each round can try multiple keys; the total timeout_s budget also includes queueing and retry waits. Capacity is released while waiting between attempts. Exhausted account quota, rejected declarations, and responses without an observed hosted call do not trigger another retry round; another capable target may still be tried. Failure messages distinguish these cases and suggest an appropriate next action.
Responses remote MCP errors fail the call. Requests requiring provider-side approval stop and direct you to the local MCP integration for interactive approval; the bridge never automatically approves them. Native message citations, file references, and unknown output fields are retained. Full native output and truncated call payloads are saved as session artifacts with readable references in the tool result. Provider files currently retain container_id, file_id, and filename references; Chord does not automatically download these files. Configure image_paths to attach base64 image data actually returned by the provider.
The bridge currently supports messages and responses sub-requests. Main conversation history uses ordinary tool results. Hosted declarations for Gemini and Chat Completions, and native main-conversation replay, require separate protocol adapters.
Project memory (automatic extraction)
Section titled “Project memory (automatic extraction)”The top-level memory section controls automatic cross-session memory extraction. Reading an existing project MEMORY.md is always automatic and needs no config; this key only decides whether Chord sends frozen history sessions to the model to grow memory records and writes project files. What gets stored, how the summary loads, and how to review or remove entries: Project Memory.
memory: enabled: true model_pool: memory-extract| Field | Default | Description |
|---|---|---|
enabled |
false |
Enable automatic memory extraction for this machine + project. When on, frozen sessions may be sent to the model and auto-written into MEMORY.md / .chord/memory/records/ as ordinary project files. When off (or unset), Chord never sends history to the model and never writes memory files, but still loads an existing MEMORY.md. |
model_pool |
(unset) | Name of a model_pools entry used for extraction requests instead of the main model pool. Use it to choose extraction models and reasoning settings independently of the main conversation. When unset, extraction uses the main model pool. Both paths preserve each model’s configured reasoning settings; an omitted effort is not overridden and retains the provider default. A configured pool must be defined in model_pools; otherwise extraction stops with a setup failure naming the missing pool. |
Precedence
Section titled “Precedence”- May appear in the global config and in project
.chord/config.yaml; project values override the user-level value like every other setting. - Because a project can enable extraction for itself, opening a project with
memory.enabled: truemay start uploading that project’s history sessions to the model. When enabled, the status bar shows aMEMORYindicator so the state is visible; a stalled extraction turns it intoMEMORY-FAILand adds a one-line notice with the reason. - The value is read at startup; changing the config file requires a restart of the running process.
Multi-agent orchestration resource limits
Section titled “Multi-agent orchestration resource limits”The top-level orchestration section bounds process-local resources used by MainAgent/SubAgent workflows. It does not grant tool permissions or change per-agent delegation limits such as delegation.max_children; it limits how many admitted runtimes and LLM requests can run at once, how much SubAgent input can queue, and how much mailbox data remains in memory.
Most users should keep the built-in defaults. Configure these limits when a provider has a strict concurrency quota, the host has limited memory, or orchestration metrics show sustained queueing or rejection.
orchestration: max_live_runtimes: 10 max_borrowed_runtimes: 1 max_bypass_runtimes: 4 max_active_llm_requests: 10 provider_max_active_requests: openai: 6 anthropic: 4 model_max_active_requests: openai/gpt-5.5: 3 subagent_queue_messages: 256 subagent_queue_bytes: 4194304 # 4 MiB mailbox_memory_messages: 512 mailbox_memory_bytes: 8388608 # 8 MiB subagent_compact_usage: 0.8 waiting_main_expiry_turns: 5 waiting_main_min_wait_sec: 300 # 5 minutes waiting_main_max_wait_sec: 3600 # 1 hour| Field | Default | Description |
|---|---|---|
max_live_runtimes |
10 |
Maximum normally admitted Agent runtimes. Further normal runtime acquisition waits until a slot is released. Wake reactivation may use the separately bounded borrowed or bypass pools when ordinary capacity is exhausted. |
max_borrowed_runtimes |
1 |
Additional temporary runtime admissions used to wake orchestration work that must make progress, such as a parent resuming after a child event. Borrowing is bounded separately from normal runtime slots. |
max_bypass_runtimes |
4 |
Maximum wake reactivations that may bypass both the normal and borrowed runtime pools when neither can make progress. When it is exhausted, the wake is refused and the durable message remains queued. |
max_active_llm_requests |
10 |
Process-wide maximum concurrent LLM requests across orchestrated agents. Eligible requests wait when the limit is full. |
provider_max_active_requests |
none | Optional concurrent-request limits keyed by provider name, for example openai. A request must satisfy this limit and the process-wide limit. |
model_max_active_requests |
none | Optional concurrent-request limits keyed by provider/model. Inline variants such as @high are ignored for matching, so openai/gpt-5.5 covers all variants of that model. |
subagent_queue_messages |
256 |
Maximum pending input messages for each SubAgent. A new enqueue is rejected when either this count or the byte limit is reached; existing queued messages are preserved. |
subagent_queue_bytes |
4194304 |
Maximum estimated bytes of pending input for each SubAgent. This is an in-memory admission bound, not a disk spool. |
mailbox_memory_messages |
512 |
Maximum SubAgent mailbox messages retained in memory across the MainAgent inbox and owner-specific mailboxes. |
mailbox_memory_bytes |
8388608 |
Maximum estimated bytes retained by those in-memory mailboxes. Durable non-progress messages that exceed the memory budget are referenced through the on-disk mailbox spool; progress updates may be coalesced or omitted from memory. |
subagent_compact_usage |
0.8 |
Proactively compress a SubAgent’s context when estimated usage reaches this fraction of its usable input budget. The default matches context.compaction.threshold; SubAgents use local token estimates and a lightweight sliding-window checkpoint rather than MainAgent’s usage-driven compaction pipeline. Must be greater than 0 and less than 1. |
waiting_main_expiry_turns |
5 |
User-turn budget for a SubAgent parked while waiting for its owner. The turn budget expires only after waiting_main_min_wait_sec has also elapsed; waiting_main_max_wait_sec still expires the wait unconditionally. |
waiting_main_min_wait_sec |
300 |
Minimum wall-clock wait, in seconds, before the turn budget can expire a waiting_main task. |
waiting_main_max_wait_sec |
3600 |
Maximum wall-clock wait, in seconds, after which a waiting_main task expires regardless of user-turn activity. The effective value is never below waiting_main_min_wait_sec; when both clocks are set explicitly and this maximum is below the minimum, loading fails instead of clamping. |
Precedence and value rules
Section titled “Precedence and value rules”- These settings may appear in the global config and in project
.chord/config.yaml. Positive project scalar values override the corresponding global values. provider_max_active_requestsandmodel_max_active_requestsare merged by key. A project entry replaces the same global key while preserving unrelated global entries.- Scalar values that are zero or negative do not mean “unlimited”: they retain the inherited or built-in default.
subagent_compact_usageis only valid strictly between0and1: an out-of-range value (including0) is ignored with a warning, a project value then inherits the merged global value, and an unset global falls back to0.8. Unlikecontext.compaction.threshold: 0, zero does not disable SubAgent context protection. - Only positive provider/model map limits are enforced. Keep map keys explicit and use positive integers; do not rely on zero as a general unlimited-mode switch.
- Limits are process-local. They do not coordinate quotas across multiple Chord processes.
Tuning guidance
Section titled “Tuning guidance”- To comply with an API quota, set the provider or model limit first; keep
max_active_llm_requestsas the overall safety ceiling. - Keep
max_bypass_runtimessmall and positive. It exists only to let wake reactivations make progress when normal and borrowed capacity are exhausted; it is not ordinary throughput capacity. - On a memory-constrained host, reduce mailbox byte/message limits gradually. Overflow uses durable storage, so lower limits trade memory for additional disk I/O.
- Reduce SubAgent queue limits only when producers can handle enqueue rejection. These queues do not spill to disk, and overly small limits can interrupt parent/child coordination.
- Keep
max_borrowed_runtimessmall but positive. Borrowed slots exist to break orchestration progress stalls, not to increase ordinary throughput. - A
waiting_maintask expires when its turn budget and minimum wait are both satisfied, or when the maximum wait is reached. Increase the turn budget or minimum wait when owners need more time to respond; increase the maximum only when parked tasks should remain recoverable for longer. - Lowering
subagent_compact_usagereduces context-overflow risk but causes earlier and more frequent compression. Raising it reduces compression work but leaves less recovery headroom. - Increasing concurrency is not automatically faster: provider throttling, model latency, local memory pressure, and workspace lease contention can reduce effective throughput. Change limits using observed queue/rejection metrics and end-to-end latency rather than CPU count alone.
MCP servers connect in two ways: Chord launches a local command and exchanges JSON-RPC over stdio, or it connects to a remote HTTP endpoint.
Local command (stdio)
Section titled “Local command (stdio)”mcp: chrome-devtools: command: "npx" args: ["-y", "chrome-devtools-mcp@latest"]command is the executable to launch, args are its arguments, and optional env entries are appended to the inherited environment. Chord starts the process and talks to it over stdin/stdout using newline-delimited JSON-RPC.
HTTP server
Section titled “HTTP server”MCP servers can expose many tools. Use allowed_tools to expose only selected remote tool names and avoid sending unused tool schemas to the model:
mcp: search: url: https://mcp.exa.ai/mcp allowed_tools: - web_search_exa - web_fetch_exaThe server name (search above) is user-defined. With this example, Chord registers only mcp_search_web_search_exa and mcp_search_web_fetch_exa. Filtered tools are not registered and do not enter the LLM tool surface.
HTTP request headers
Section titled “HTTP request headers”For remote MCP servers that require authentication, use headers to send extra headers with every request. A service like Exa requires an x-api-key:
mcp: exa: url: https://mcp.exa.ai/mcp headers: x-api-key: "$EXA_API_KEY"A header value starting with $ is expanded from the environment (here EXA_API_KEY), so secrets do not have to be written into the config file; a $ value that expands to an empty string is a configuration error, since it would authenticate with a blank credential. Header names must be valid HTTP header names, and values must not contain CR or LF.
headers applies only to remote (url) servers; stdio servers do not carry HTTP requests, and configuring headers for one is rejected. Protocol-managed headers (Content-Type, Accept, Mcp-Session-Id) are set by Chord and are not affected by headers.
Manual (on-demand) MCP servers
Section titled “Manual (on-demand) MCP servers”By default, configured MCP servers auto-start and become part of the default LLM tool context. For an MCP server you do not need in every conversation, set manual: true: it stays disabled at startup, Chord normally does not connect to it, and its tool descriptions are not added to the default context, reducing context overhead. Enable it manually only when you need it:
mcp: exa: url: https://mcp.exa.ai/mcp manual: true- When
manual: true, the server starts in a disabled (gray) state and does not connect until you enable it. - Only servers configured with
manual: truecan be changed at runtime with/mcp. Auto-start servers are read-only in the MCP selector and are not affected by/mcp enable|disable. - Enable/disable at runtime with
/mcp(menu in TUI) or with explicit commands:/mcp enable <server>/mcp disable <server>/mcp status
- Runtime
/mcp enable|disablechanges are allowed while a turn is running. The current in-flight request keeps the tool surface it started with; the next LLM request (including automatic retry/recovery requests) applies the new execution state. - By default, that next request rebuilds the top-level MCP tool surface and may miss the existing prompt cache. Models that explicitly enable
compat.chat_completions.mcp_system_tools_messageorcompat.responses.mcp_additional_toolsinstead receive request-only tool declarations at a fixed conversation anchor. Later requests replay each declaration at the same position; disabling a server blocks execution but keeps its existing declaration, preserving the prefix. The mount follows the selected model: a model switch rebuilds the top-level tool surface for the new model and then resumes dynamic mounts where that model accepts them. Session resume or switch and durable compaction instead revert to top-level tools for the rest of the session run, because the fixed anchors cannot be trusted; a later model switch keeps that revert, and only starting a new session run (such as/new) re-enables dynamic mounts. A tool whose earlier calls are still in the history but that has no declaration in the current run also makes the request fall back to top-level tools, so its declaration never lands after its own calls. The revert itself is silent; a later/mcp enable|disablewarns about the prompt cache again only once a request has installed the rebuilt surface. - The enabled/disabled intent for manual servers is saved with the session:
/mcp enablepersists it,/mcp disableclears it, and resuming the session (including after restart) reconnects the manual servers that were enabled when it was last active. A connection failure keeps the intent, so the server stays “enabled (unavailable)” and can be retried instead of silently returning to disabled.
Startup consistency
Section titled “Startup consistency”Auto-start MCP servers still connect asynchronously after the TUI starts, but the first LLM request waits until each auto-start server has either connected successfully or reached a terminal failure state. This avoids tool-surface inconsistency between the agent and the model.
Agent config
Section titled “Agent config”Built-in roles include builder and planner. Both are main-mode, so
delegate is not registered until you define at least one mode: subagent
role of your own; see mode below. You can also add custom
agents or override built-ins. Agent files can live in:
~/.config/chord/agents/.chord/agents/
Supported file formats:
.md: YAML frontmatter plus a Markdown body. The body becomes the system prompt..yaml/.yml: plain YAML. Usepromptorsystem_promptfor the system prompt.
Markdown agent example:
---name: backend-coderdescription: Backend developermode: subagentpermission: write: ask edit: ask---
You are an agent focused on backend development.Equivalent YAML agent example:
name: backend-coderdescription: Backend developermode: subagentpermission: write: ask edit: askprompt: | You are an agent focused on backend development.Common fields include:
name: agent name. If omitted, Chord uses the filename without extension. If specified, it must match the filename without extension (for example,builder.yamlmust declarename: builder). A single directory cannot contain duplicate agent names, including duplicates across.md,.yaml, and.yml. Project-level agents may still override same-named global agents by design.description: short description shown to the main agent when delegation is available. Put routing intent here; Chord does not keep a separate annotation layer for preferred tasks or write mode. Whether a role may write files is decided bypermission, and the Delegate tool surfaces that asempty_scope=allowedornon_empty_scope=requiredon each agent choice.mode:mainfor a MainAgent role, orsubagentfor a SubAgent. Empty and unknown values behave asmain;sub_agentandsubare accepted as SubAgent aliases. Thedelegatetool is registered only when at least onesubagentrole is visible to the delegating role, so a configuration with no subagent definitions has no delegation surface at all: that is the usual reasondelegateappears to be missing.model_pools: optional ordered list of pool names this agent can use. Pool definitions live inconfig.yamltop-levelmodel_pools; when omitted, the agent can use all top-level pools sorted by name. Inline variants such asopenai/gpt-5.5@highare specified in the pool definitions.variant: default variant when a model ref does not include@variant.permission: per-tool permission policy for this agent. Permissions live directly in agent config files; when the confirmation popup remembers a rule,projectupdates the current project’s.chord/agents/<role>.yaml, andglobalupdates the user config directory’sagents/<role>.yaml(default:~/.config/chord/agents/<role>.yaml). Chord no longer writes a separate permissions directory. Some orchestration tools have special semantics (delegatepatterns matchagent_typeand also gate delegated-work controls such ascancel;handoffanddonetreatallowandaskas workflow-available states with Chord’s own confirmation gates). See Permissions & Safety before relying on fine-grained control-tool rules.mcp: additional auto-start MCP servers scoped to this agent. Agent MCP is additive: a server name already present in the effective global/projectmcpconfig is a startup error. Agent-scoped servers cannot usemanual: truebecause runtime MCP controls manage the top-level server surface; configure a manual server at the project/global level instead. Remove an agent entry to inherit a top-level server, rename it for a separate private server, or override the top-level server in.chord/config.yamlfor the whole project. Different agents may reuse the same private server name without sharing the connection unless they are instances of the same agent definition.delegation: delegation limits for this agent definition. Values above the ceiling or negative values are configuration errors:max_children: how many direct, still-active child tasks this agent may have at once. Defaults to10; the ceiling is64.max_depth: how deep nested delegation may go. It is evaluated per worker, against the worker’s own definition: a SubAgent’s ability to delegate further is checked against its owndelegation.max_depthand its current depth, never against its parent’s or the root role’s setting, so a root role withmax_depth: 1cannot stop a child definition that declaresmax_depth: 8from nesting deeper. Defaults to1(a first-level SubAgent cannot delegate further until its own definition raises the value); the ceiling is8.child_join: whether children a SubAgent delegates stay tied to the owner’s task. Defaults totrue: the owner cannot complete while joined children are still running, so its completion is deferred until they finish or are explicitly stopped, and a cancelled or failed owner cancels its joined children with it. Withfalse, the owner may finish early and its still-running children detach and continue under the main agent instead of being cancelled. Only nested delegation is affected: children delegated by the main agent never join, because the main agent is not itself a task.
prompt/system_prompt: system prompt for plain YAML files. Setting either one replaces any built-in prompt block the role would otherwise get.prompt_preset: selects a built-in role prompt block by capability instead of by role name. Accepted values areplanningandnone.planninginjects the built-in planning block (plan-document naming and format, the direct-answer-versus-plan decision, handoff ordering, and plan quality rules); it also suppresses the bug-triage block, which would otherwise duplicate the planning workflow’s own investigation outline.nonesuppresses any built-in block. When the field is omitted, the role gets no built-in block whatever it is called: the role name never selects one, so a custom role namedplannerhas to declareprompt_preset: planningto keep the planning block. Unknown values are a configuration error.prompt_append: text appended after the effective role prompt: after the preset block, or afterprompt/system_promptwhen the role replaces it. Use this to add project conventions without taking over maintenance of the whole block, which also keeps the preset’s tool-aware wording (it adapts to the tools the role can actually see).
A custom planning role that reuses the built-in block:
name: architectdescription: Architecture planning rolemode: mainprompt_preset: planningprompt_append: | Reference the relevant ADR number in every plan document.permission: "*": deny read: allow grep: allow glob: allow write: .chord/plans/*: allow handoff: allowExample:
name: buildermode: mainmodel_pools: [default]permission: "*": deny read: allow view_image: allow grep: allow glob: allow web_fetch: "localhost:8000": ask shell: allow edit: ask write: askIn an allowlist like this, the leading "*": deny covers every tool you did not
list, so anything the role needs must be named. Two exceptions are worth
knowing, because they would otherwise look like the feature is broken:
compact_contextanddoneare not covered by the wildcard. The switch that makes each reachable (context.compaction.model_drivenfor the former, starting a loop for the latter) is itself the authorization, so this role can run model-driven compaction and loop mode without listing them. Name a tool explicitly (done: deny) when you do want to withhold it. See Permissions & Safety.- Everything else is covered normally.
todo_writeandquestionare ordinary tools here: leaving them out means the model tracks no TODO list and asks the user in plain assistant text instead of a structured prompt. Both degrade cleanly, so add them only if you want those channels.
Context management
Section titled “Context management”Long-session context handling covers context compaction (LLM-generated
summaries that rewrite session history) and context reduction (request-time
trimming of stale tool output). Both are configured under the top-level
context: key and documented on their own page:
Context management.
Post-tool diagnostics
Section titled “Post-tool diagnostics”After edit, apply_patch, or write modifies a file, Chord can append language diagnostics to the tool result so the model sees compile or lint problems immediately. This is controlled by the diagnostics config and is enabled by default for Python (an LSP semantic backend with a Ruff quick fallback). Set diagnostics.enabled: false to skip the whole pipeline.
Native file tools also send workspace/didChangeWatchedFiles events to matching LSP servers before syncing the textDocument: write sends Created for new files, write on existing files plus edit / apply_patch send Changed, and successful delete sends Deleted. This helps Pyright, TypeScript, gopls, rust-analyzer, and similar servers refresh their project graph promptly, reducing transient unresolved-import/module diagnostics after new files are created.
Diagnostics are still returned immediately in file-tool results so the model can attribute problems to the current edit; files created or removed by shell commands or external programs are not yet reported through a full filesystem watcher.
For Python, two backends are used:
diagnostics.python.semantic_backend: the primary LSP server (defaultpyright). Itsserverfield must match a server key underlspso the language server is actually configured.diagnostics.python.quick_backend: a one-shot fallback (defaultruff check) used for large files, or when the semantic backend is unavailable.
diagnostics.python.large_file.{line_threshold, byte_threshold, strategy} decides when a file is large enough to use the quick backend instead of the semantic one; run_semantic_when_quick_unavailable: true forces the semantic backend even on large files when the quick backend is missing. Ruff quick diagnostics do not update the LSP sidebar: they appear only in edit, apply_patch, or write results and note that full semantic diagnostics were skipped.
Recommended Python skeleton:
lsp: pyright: command: pyright-langserver args: ["--stdio"] file_types: [".py", ".pyi"]
diagnostics: python: semantic_backend: server: pyright quick_backend: type: command command: ruffdiagnostics.python.output.{max_near_diagnostics, max_outside_diagnostics, max_total_diagnostics, near_range_before_lines, near_range_after_lines} shapes how much appended diagnostics text is shown, prioritizing errors and warnings before info and hints. See the Configuration cheatsheet for the full field list.
Diagnostics appended to the tool result cover the edited files’ own problems, plus cached problems from other files in the same directory as an edited file (Go packages are compiled per directory, and workspace diagnostics cover every file in a diagnosed package). Each other-file diagnostic is attached only once per session: the same problem is not repeated in later tool results until that diagnostic disappears from the server’s published set, after which a reappearing problem is reported again.
Later edits do not repeat it either: a problem the model already has costs context to restate, so a surviving diagnostic stays suppressed and only changes are reported. A problem that was fixed (by another agent, a file copy, or a git checkout restore) stops being reported instead of being served from cache: the cached diagnostics are withheld as soon as the file no longer matches what they were computed from, and they are dropped once the server publishes without them.
Resuming a session keeps the suppression instead of restarting it: diagnostics already rendered in the restored transcript are recovered from it, so --continue does not re-announce problems that are already visible earlier in the same conversation.
Diagnostics for a file that changed on disk since the server last published them (for example, fixed by another editor or process) are skipped until Chord synchronizes the file and receives fresh diagnostics, because the cached result may no longer reflect its current content.
Provider/model diagnostics
Section titled “Provider/model diagnostics”# smoke-test all providers with representative modelschord doctor models
# test one provider's representative modelchord doctor models --provider openai
# test an exact model or variantchord doctor models --model openai/gpt-5.5@highchord doctor models --provider openai --model gpt-5.5@high
# audit each entry in a model pool independentlychord doctor models --pool thinkingUse this command as an auth, endpoint, transport, model, and variant tuning smoke test. It uses the same merged global + project config view as normal runtime startup, so project-level provider/proxy/model overrides are included. Pool diagnostics request each pool entry independently rather than following the normal fallback chain.
Configuration cheatsheet
Section titled “Configuration cheatsheet”The full top-level keys of config.yaml (both global ~/.config/chord/config.yaml and project-level .chord/config.yaml). All keys are optional unless noted.
| Key | Type | Default | Scope | Summary |
|---|---|---|---|---|
providers |
map[name]Provider |
— | global / project | Per-provider config (type, api_url, preset, key_rotation, key_order, models, compress). See Minimal provider config. |
model_templates |
map[name]YAML |
empty | global / project | YAML-anchor namespace only; entries are reusable through aliases and are not runtime model definitions. |
model_pools |
map[name][]ref |
— | global / project | Reusable named pools of full provider/model[@variant] refs. See Model pools. |
thinking_translation |
object | disabled (max_chars: 1000) |
global / project | Optional appended translation preview for thinking / reasoning cards. Requires target_language and model_pool; failures only skip the affected thinking block. |
context |
object | see below | global / project | compaction and reduction settings. See Context compaction and Context reduction. |
diagnostics |
object | enabled (Python LSP + Ruff fallback) | global / project | Post-tool diagnostics appended to edit, apply_patch, or write results. diagnostics.python.semantic_backend is the primary LSP server (default pyright); diagnostics.python.quick_backend is a one-shot fallback (default ruff check). diagnostics.python.large_file.{line_threshold, byte_threshold, strategy} controls when large files use the quick backend, and run_semantic_when_quick_unavailable: true forces semantic diagnostics when the quick backend is missing. diagnostics.python.output.{max_near_diagnostics, max_outside_diagnostics, max_total_diagnostics, near_range_before_lines, near_range_after_lines} shapes the appended diagnostics text. Diagnostics are shown by severity priority (errors/warnings first, then info/hints if slots remain). Set diagnostics.enabled: false to skip the whole pipeline. |
skills |
object | empty | global / project | paths: [...] — additional skill directories beyond the defaults. |
confirm_timeout |
int (seconds) | 0 (no timeout) |
global / project | Timeout for confirmation dialogs in TUI; 0 means wait forever. |
question_timeout |
int (seconds) | 0 (no timeout) |
global / project | Timeout for the Question tool in TUI and headless; 0 means wait forever. The countdown covers the whole wait, including time queued behind another dialog, and each question in a batch is timed separately. When it elapses, the question closes as no_response; Chord never adopts an answer automatically. |
diff |
object | {inline_max_columns: 200} |
global / project | TUI diff rendering. inline_max_columns caps one-line inline diff width. |
desktop_notification |
bool | false |
global / project | Enable local-TUI terminal notifications; Chord auto-selects OSC 9 or OSC 777 by terminal and pairs each notification with a terminal bell (BEL). (Unsupported terminals ignore the sequence; see Platforms.) |
desktop_notification_foreground |
bool | true |
global / project | Send local-TUI terminal notifications (escape sequence and bell) while the terminal is focused. Set to false for background-only notifications. |
prevent_sleep |
bool | false |
global / project | Prevent macOS idle sleep while any agent is active. macOS-only; no-op elsewhere. |
keymap |
map[action][]key |
see Keybindings | global / project | Override key bindings. Action names use lower snake_case. |
commands |
map[/cmd]text |
empty | global / project | Custom slash commands; "/cmd" → text inserted as a user message. See Customization: Custom slash commands. |
ime_switch_target |
string | empty | global / project | IM identifier passed to im-select / im-select.exe when entering Normal mode. Linux/macOS/Windows. |
log_level |
string | info |
global / project | debug / info / warn / error. debug is verbose. |
paths |
object | XDG defaults | global only | state_dir, cache_dir, sessions_dir, logs_dir. CLI flags and CHORD_* env vars override. |
maintenance |
object | disabled | global only | size_check_on_startup, warn_state_bytes, warn_cache_bytes. |
lsp |
map[name]Server |
empty | global / project | Per-language-server config. See Customization: LSP. |
mcp |
map[name]MCP |
empty | global / project / agent | Per-MCP-server config. See MCP. |
hooks |
object | empty | global / project / agent | Hooks per trigger point. See Hooks. |
max_output_tokens |
int | 64000 |
global / project | Global cap on requested output tokens. Effective limit is also clamped by each model’s limit.output; reasoning requests also respect it. |
stream_retry_rounds |
int | 0 (retry until success/cancel) |
global / project | Hard cap on public LLM full-round retries. 0 keeps retrying until success, cancellation, or terminal failure. |
proxy |
string | empty (use env / direct) | global / project | Global proxy URL. Per-tool override via web_fetch.proxy. |
web_fetch |
object | empty | global / project | user_agent, proxy (inherits global if nil; empty string = direct). See WebFetch. |
worktree |
object | empty | global / project | branch_prefix names worktree branches; root sets where checkouts are created (a relative value resolves against the repository root; the default stays outside it under the state dir). See Worktrees. |
Provider field reference
Section titled “Provider field reference”Chord automatically propagates the current Chord session id to OpenAI-family providers as cache/routing affinity metadata: OpenAI Responses requests include prompt_cache_key, and OpenAI Chat Completions / Responses HTTP requests include X-Session-Id and session-id headers when a session id is available.
The key is per client rather than per provider: the main agent uses the current Chord session id, and each SubAgent derives its own <session>:sub:<instanceID> key so one agent’s requests never inherit another’s cache identity. These fields are not user-configurable; they follow the active Chord session and are cleared or changed on session switch/resume.
Anthropic prompt caching is driven by cache_control blocks, and Chord also sends JSON-formatted metadata.user_id automatically with a stable anonymous device_id plus a stable routing session_id derived from local/provider identity. These Anthropic metadata fields are not user-configurable.
In explicit mode (the default for Anthropic models), Chord places up to four cache_control breakpoints by priority: the last system block, the frozen reduced-prefix boundary (when incremental reduction has frozen a stable prefix), the newest durable message, and the last assistant message, so long agent loops reuse the frozen historical surface instead of re-writing the moving tail each turn.
The newest breakpoint deliberately skips request-scoped overlays (runtime hints appended after the conversation tail), because those bytes are gone on the next request and a cache entry written past them could never be read back.
For Anthropic models, prompt_cache.ttl accepts 5m (the default when
omitted) and 1h, and applies to every breakpoint Chord places in both auto
and explicit mode:
providers: anthropic: models: claude-sonnet-4-5: prompt_cache: ttl: 1hGemini does not have a simple per-request session-id cache key in Chord’s
generateContent transport; its cache signals come from provider-specific
cached-content APIs/usage fields, not from a Chord session id header.
| Field | Type | Description |
|---|---|---|
type |
string | messages / chat-completions / responses / generate-content. Auto-detected from api_url or preset when omitted. |
api_url |
string | Endpoint URL. Chord detects provider type from the URL path, ignoring query strings and fragments. For Gemini, the /models base path; Chord appends /{model}:streamGenerateContent?alt=sse. For Azure Responses, ?api-version=... is optional and can be used to pin a specific API version. |
preset |
string | codex (OpenAI Codex / ChatGPT OAuth). Azure OpenAI Responses uses a plain type: responses provider with auth_scheme: api-key, store: true, and compat.request_overrides.headers set to null for the Codex identity headers. |
trust_http_400 |
bool | Treat HTTP 400 as a terminal request error. preset: codex defaults to true; aggregating/proxy gateways default to false because they often wrap upstream overload as 400. |
retry_after_max_s |
int | Longest Retry-After wait honored, in seconds (1-86400). The header always applies as the key cooldown, ahead of retry_backoff/retry_delay_ms; this only bounds how long a single hint may block a key. preset: codex defaults to 86400; third-party gateways, which can echo arbitrary values, default to 60. |
key_rotation |
string | on_failure (default) / per_request. Controls when a credential / API key is reselected. |
key_order |
string | sequential (non-Codex default) / random / smart (Codex only). Controls how Chord chooses among selectable keys. |
retry_backoff |
string | exponential (default) / fixed / none. Controls generated delay between complete rounds and, when explicitly set, ordinary HTTP 429 key cooldown. Explicit settings replace Retry-After for ordinary 429s; confirmed quota resets and hard credential states still win. |
retry_delay_ms |
int | Base/fixed round and ordinary-429 delay in milliseconds, from 0 through 60000; 0 / omitted defaults to 1000ms. Setting the field—including explicit 0—is an override even when retry_backoff is omitted. Ignored for none. Out-of-range values are logged and fall back to the default instead of failing startup. |
compress |
string | Upstream request body compression encoding: gzip or zstd; unset = off. Applies only when compression shrinks the payload. The boolean compress: true form is gone — it is ignored and reported by chord doctor config (migrate to compress: gzip). |
response_header_timeout |
int | Timeout in seconds from starting a streaming HTTP request until response headers arrive, including connection setup and request-body upload. 0 / omitted uses the built-in default; healthy streams are bounded by stream_idle_timeout, not a total request timer. |
stream_idle_timeout |
int | Stream idle timeout in seconds for this provider. 0 / omitted uses built-in SSE/WebSocket idle defaults. |
stream_total_timeout |
int | Wall-clock cap in seconds for one stream, including Codex Responses WebSocket reads. 0 / omitted applies no cap — a stream that keeps producing data is never cut off by elapsed time alone. |
websocket_handshake_timeout |
int | Responses WebSocket handshake timeout in seconds. 0 / omitted uses the built-in default. |
supported_service_tiers |
list | Provider-level default accepted non-standard tiers for its models, e.g. [fast, slow] or [fast]. Model entries can override it. |
parallel_tool_calls |
bool | true — Provider-level default for Responses / Chat Completions tool parallelism; model and variant values override it. |
compat.responses.* |
object | protocol defaults — Provider-level optional Responses fields: send_store, send_reasoning_include, send_tool_choice, send_prompt_cache_key, send_max_output_tokens, and mcp_additional_tools. |
compat.responses.mcp_additional_tools |
bool | false — Mount runtime manual-MCP schemas as fixed-anchor input[type="additional_tools"] items instead of changing top-level tools. Enable only for Responses endpoints/models known to accept this item. The mount follows the currently selected target; when a fallback pool member without this capability serves the request, Chord inlines the declarations into that request’s top-level tools array. |
compat.hosted_tools |
list | (empty) — Hosted tool names this provider’s models may serve. Each call runs as a separate request that declares the tool’s per-type wire declaration and returns the provider’s result as an ordinary tool result; the main conversation request never declares it. Enable only on endpoints known to support the declaration: one that rejects it fails the call with the endpoint’s error, and one that silently ignores it fails with a message pointing back at this list. The catalog shape is documented under Hosted tools. |
compat.apply_patch.enabled |
bool | Three-state — when omitted, Chord infers from the model name. true keeps apply_patch (hiding edit, write, and delete); false falls back to edit with write/delete visible. gpt-5-and-later family names (gpt-5, gpt-5-mini, gpt-5-nano, gpt-5-codex, any gpt-5.* name, later majors like gpt-6-astra) and codex-auto-review default to true; gpt-oss-*, gpt-3.5, gpt-4/4o, the o-series, and non-OpenAI models default to false. |
compat.apply_patch.freeform |
bool | Three-state — when omitted, Chord infers from the model name and the wire type. true emits apply_patch as a freeform custom tool (type: "custom" with a grammar); false emits a JSON function tool. gpt-5-and-later family names and codex-auto-review on Responses endpoints default to true; all non-Responses wires default to false (they have no custom tool type). Hosts that accept Responses but reject custom tools have no built-in exception: set false there, or true only for gateways that actually accept custom tools. |
compat.chat_completions.send_stream_options |
bool | true — Omit stream_options.include_usage for gateways that reject it; streaming token usage then remains unavailable. |
compat.chat_completions.infer_finish_reason |
bool | false — Derive a normal stop / tool_calls completion for compatible gateways that end the stream without emitting finish_reason; otherwise those streams are treated as interrupted. |
compat.chat_completions.requires_tool_result_name |
bool | false — Emit the paired tool name on tool result messages for gateways that require name alongside tool_call_id. |
compat.chat_completions.requires_assistant_after_tool_result |
bool | false — Insert a synthetic assistant message between a tool result and the next user message for gateways that reject a user message directly after tool results. |
compat.chat_completions.mcp_system_tools_message |
bool | false — Mount runtime manual-MCP schemas as fixed-anchor role: system messages with a tools field and no content, instead of changing top-level tools. Enable only for models known to accept the Kimi-compatible dynamic-tool shape. The mount follows the currently selected target; when a fallback pool member without this capability serves the request, Chord inlines the declarations into that request’s top-level tools array. |
compat.chat_completions.keep_reasoning_effort |
bool | false — Keep reasoning_effort and reasoning request overrides active when a current-turn assistant tool-call message replays without reasoning_content. Chord otherwise reads the missing content as a backend that cannot replay reasoning and strips those controls for the rest of the turn; enable it for endpoints that accept reasoning controls without a reasoning-content replay contract, such as Grok on the Chat Completions wire. It only keeps the request-side controls in place; it does not supply reasoning content to backends whose contract validates the replayed history (DeepSeek when a request carries tools, Kimi K3, Qwen preserve_thinking). |
compat.chat_completions.native_thinking |
string | Selects the request shape Chord uses for a model’s thinking settings when the endpoint is a Chat Completions gateway that translates the call into the model’s native API. Without a value, only a DeepSeek route selects the thinking:{type} shape: a model ID that names a DeepSeek API model, or any model under compat.reasoning_continuity.contract: deepseek (contract: none drops the name shortcut). Every other model needs an explicit value. Values: gemini (extra_body.google.thinking_config), gemini-3 (the same shape with Gemini 3 signature repair), anthropic (thinking:{type,budget_tokens}), thinking (the native thinking:{type} object used by DeepSeek, GLM, Kimi K2.x, and Doubao), and qwen (enable_thinking); family names such as claude, deepseek, glm, kimi, and doubao are accepted as aliases. off (alias none) disables the conversion for endpoints that reject unknown body fields. A model that configures no thinking block never sends the field, except on a DeepSeek route: its reasoning contract enables thinking by default, while explicit thinking.type: disabled or reasoning.effort: none disables it and omits effort (see compat.reasoning_continuity.contract). The selector is also the only signal that a gateway model is Gemini or Claude: without it Chord does not write Gemini thought signatures back, and Gemini 3 rejects the request that follows each tool call (HTTP 400), so pin gemini-3 on every Gemini 3 model behind a gateway. chord doctor config warns about such models. An explicit selector governs this request shape only and always wins over the DeepSeek default; the reasoning contract stays in force, and off does not disable it. See Thinking behind a Chat Completions gateway. |
compat.usage.input_includes_cache_read |
bool | Protocol default — Override whether the provider’s top-level input count already contains cache-read tokens. Defaults: Messages false; Chat Completions / Responses / Generate Content true. |
compat.usage.input_includes_cache_write |
bool | Protocol default — Override whether the provider’s top-level input count already contains cache-write/cache-creation tokens. Defaults: Chat Completions / Responses true; Messages / Generate Content false. |
models |
map | Map of model id → model config. |
Model field reference
Section titled “Model field reference”| Field | Type | Description |
|---|---|---|
limit.context |
int | Total request window in tokens when known. If limit.input is omitted, Chord derives the input budget from this minus the model’s limit.output (falling back to the max_output_tokens default when no output cap is declared). |
limit.input |
int | Separate input cap when a provider publishes one. Chord uses it to compact or retry before the prompt is too large. |
limit.output |
int | Maximum output tokens; runtime is also clamped by max_output_tokens. |
compaction |
object | Per-model compaction overrides: compaction.threshold (auto-compaction usage ratio; 0 disables for this model) and compaction.reminder (pressure-reminder line; derived from threshold when absent, -1 disables the reminder only). Unset fields inherit the global context.compaction.*. Out-of-range values are rejected with a warning and inherit the global value. Derivation and tuning guidance: Context compaction. |
reasoning |
object | OpenAI reasoning options. reasoning.effort passes through without a local whitelist, so any provider-supported level (e.g. GLM max / minimal / none) reaches the upstream unchanged; Responses normalizes whitespace and casing before sending (unset = omit and use provider/model default). For Responses, reasoning.summary supports auto / concise / detailed / none; when reasoning is active, unset defaults to auto, while none opts out explicitly. |
text.verbosity |
string | Optional OpenAI text verbosity hint where supported; leave unset to use the provider/model default unless you intentionally want low / medium / high. |
thinking |
object | Extended-thinking options. Messages: type: adaptive carries no token budget and pairs with thinking.effort, which Chord sends as output_config.effort; Claude type: enabled requires thinking.budget; DeepSeek instead accepts type: enabled with thinking.effort and no budget; display applies only to enabled / adaptive. Gemini: thinking.level / thinking.budget / thinking.include_thoughts map into the generation request (see Google Gemini). |
compat.reasoning_continuity.mode |
string | Optional continuity override. Use openai_visible for Chat Completions models that require unchanged assistant reasoning_content and can accept portable visible reasoning from other wires; it also enables the missing-reasoning_text fallback for Responses targets with that continuity contract. Use anthropic_unsigned only for verified Messages-compatible models that replay or accept visible unsigned thinking; use none to opt out of a provider-level default. DeepSeek Chat/Messages targets ignore this field, including none: their reasoning contract fixes the mode (openai_visible on Chat, anthropic_unsigned on Messages). |
compat.reasoning_continuity.contract |
string | Declares an endpoint-specific request contract, overriding what the model name implies. deepseek selects the DeepSeek tool-history passback rules and request tuning on the Chat Completions and Messages wires (on Chat Completions also the thinking:{type} shape while native_thinking is unset), which an alias whose model ID does not identify the backend needs; gemini-3 enables missing thought-signature repair on the native Gemini endpoint (behind a Chat Completions gateway, native_thinking: gemini-3 does this instead); none opts a route out of whatever its model name implies, including that shape, for example a third-party route whose model ID names a DeepSeek API model but serves another backend, on the Chat Completions and Messages wires alike. A model ID that names a DeepSeek API model selects deepseek on its own, and on the native Gemini endpoint a model ID whose final component starts with gemini-3 selects gemini-3; other endpoints keep the generic continuity behavior. An unknown value is rejected by the config loader. |
compat.reasoning_continuity.reasoning_replay |
string | How much reasoning from completed turns is replayed. DeepSeek Chat/Messages always replay the complete history regardless of this window. current_turn (default) keeps only reasoning after the last user message; all replays completed turns unchanged for backends whose contract requires the full assistant history (Kimi K3 / keep: all, Qwen preserve_thinking, GLM clear_thinking: false, Xiaomi MiMo); none strips reasoning everywhere including the current turn, for endpoints verified to accept that. |
compat.forced_tool_choice.suppress_in_thinking |
bool | Downgrade loop-forced tool_choice: required to the backend default while reasoning/thinking is active, for OpenAI-compatible endpoints that reject forced tool choice in thinking mode. |
compat.forced_tool_choice.auto_only |
bool | Downgrade any non-auto tool_choice to the backend default unconditionally, for backends that only support tool_choice: "auto". Chat Completions / Messages / Gemini omit the field; Responses still sends "auto" unless compat.responses.send_tool_choice is false. |
compat.request_overrides.body |
object | Recursive JSON patch applied after Chord constructs the protocol request. null deletes a field. |
compat.request_overrides.rename_body_fields |
map | Renames final JSON fields while preserving Chord’s computed values. A null target deletes the source field. |
compat.request_overrides.headers |
map | Sets final request headers. A null value removes that header. |
compat.chat_completions.mcp_system_tools_message |
bool | Model-level override for the provider default described above. |
compat.chat_completions.keep_reasoning_effort |
bool | Model-level override for the provider default described above. |
compat.chat_completions.native_thinking |
string | Model-level override for the provider default described above. |
compat.responses.mcp_additional_tools |
bool | Model-level override for the provider default described above. |
compat.hosted_tools |
list | Model-level list that replaces the provider default described above. |
compat.apply_patch.enabled |
bool | Model-level override for the provider default described above. |
compat.apply_patch.freeform |
bool | Model-level override for the provider default described above. |
variants |
map | Named parameter presets. Reference with provider/model@variant. |
modalities.input |
array | Subset of text / image / pdf. Defaults to [text]; declare image / pdf explicitly when supported. |
supported_service_tiers |
list | Provider-level default or model-level override for accepted non-standard tiers, e.g. [fast, slow] or [fast]. Omit to use preset defaults. |
Service tiers and prompt caching are provider-specific. OpenAI supports priority/flex-style tiering plus prompt_cache_key / prompt_cache_retention; Anthropic supports cache_control with 5m and 1h TTLs and service-tier controls; Gemini uses its own routing / thinking / cached-content mechanisms when available. Chord maps the user-facing tier to the closest supported provider behavior instead of forcing one wire format across all backends.
For a compatible gateway whose usage fields differ from its declared protocol, set the usage semantics explicitly. For example, a Responses-compatible gateway that reports uncached input separately from both cache buckets needs:
compat: usage: input_includes_cache_read: false input_includes_cache_write: falseA Messages-compatible gateway that reports inclusive usage (input_tokens is
the full input including cache hits, and cache_read_input_tokens is only the
hit subset) needs the opposite override. Otherwise, Chord counts the cache
reads twice and understates the cache-hit rate:
compat: usage: input_includes_cache_read: trueOpenAI reasoning items can also be returned as reasoning.encrypted_content when you need stateless continuation. Treat that field as opaque continuation data: it is not meant to be rendered directly in the UI. When a readable summary is available, that is the user-facing form to show.