返回 CodeWhale
PROVIDERS.md
根目录 / docs / PROVIDERS.md
1 # Provider Registry
2
3 > 阅读简体中文版:[zh_hans/PROVIDERS.md](zh_hans/PROVIDERS.md)
4
5 This registry describes provider behavior that is wired into the current
6 Codewhale codebase. It is intentionally conservative: shipped entries are
7 limited to provider IDs, config keys, auth paths, base URLs, model resolution,
8 and capability metadata that the code already knows about.
9
10 DeepSeek remains the default provider, but every entry in `ProviderKind::ALL`
11 is a first-class selectable provider route. `ALL` is the catalog/picker
12 surface — one identity per vendor. Dual-wire dialect kinds (`*Anthropic`, e.g.
13 `deepseek-anthropic`) and the Model Studio plan variants stay on the enum for
14 serde and `provider_for_kind` but are deliberately **not** catalog rows:
15 a plan is `mode`/`base_url` and a dialect is `wire = openai|anthropic` on the
16 primary provider config (`crates/config/src/provider_kind.rs:221-226`). Hosted
17 routes, generic OpenAI-compatible endpoints, the OpenAI Codex/ChatGPT route,
18 native Anthropic, and local runtimes all run the same terminal harness against
19 the selected provider/model/base URL.
20
21 A host reached over plain Chat Completions is an ordinary named provider,
22 not a `ProviderKind`: enum variants are reserved for distinct *wires*
23 (Anthropic Messages, Codex Responses, Google thought signatures). Any such
24 host is a `[providers.<name>]` table with a base URL, a model, and a key env
25 (`docs/CONFIGURATION.md`); `/provider` and `/setup` keep a "paste a Base URL
26 and a key" path for exactly this. Offerings come from live `GET /v1/models`
27 plus the Codewhale catalog rather than a compiled roster (#5350, #6289).
28
29 Known-good hosts. These ship as bundled descriptor rows in
30 `crates/config/assets/provider_descriptors.json` (compiled in by
31 `crates/config/src/descriptors.rs`): each row says how to reach the host — wire,
32 base URL, key env, aliases — while model ids stay live from `GET /v1/models`
33 and the Codewhale catalog; the example model is only a bootstrap hint. Verify
34 against the vendor's own docs before trusting any value here:
35
36 | Host | Base URL | Example models | API key env |
37 | --- | --- | --- | --- |
38 | SenseNova | `https://token.sensenova.cn/v1` | `deepseek-v4-flash` | `SENSENOVA_API_KEY` |
39 | Baseten | `https://inference.baseten.co/v1` | `deepseek-ai/DeepSeek-V4-Pro` | `BASETEN_API_KEY` |
40 | Groq | `https://api.groq.com/openai/v1` | `llama-3.3-70b-versatile` | `GROQ_API_KEY` |
41 | Cerebras | `https://api.cerebras.ai/v1` | `llama-3.3-70b` | `CEREBRAS_API_KEY` |
42 | Command Code | `https://api.commandcode.ai/provider/v1` | `deepseek/deepseek-v4-flash` | `COMMAND_CODE_API_KEY` |
43 | Alibaba Model Studio (DashScope) | `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` | `qwen3.8-flash` | `DASHSCOPE_API_KEY` |
44 | AICraft | `https://aicraftapi.com/v1` | `claude-4.6-sonnet`; DeepSeek / Claude / Gemini / Qwen / GLM / MiniMax / Doubao families | `AICRAFT_API_KEY` |
45 | Tsubasa | `https://api.tsubasa.sh/v1` | `tsubasa-pro`, `tsubasa-fast` (32,768-token context) | `TSUBASA_API_KEY` |
46 | Cheaper Inference | `https://api.cheaperinference.com/v1` | `gpt-5.4-mini`, `claude-sonnet-5`, `gemini-3.1-pro` | `CHEAPER_INFERENCE_API_KEY` |
47 | Yolo-Auto | `https://yolo-auto.com/v1` | `qwen3.8-flash`; `yolo` / `yolo-small` | `YOLO_AUTO_API_KEY` |
48
49 AICraft's roster spans DeepSeek, Anthropic Claude, Google Gemini, Qwen, GLM,
50 MiniMax and Doubao ids on its OpenAI-compatible endpoint. The authority is
51 `GET https://aicraftapi.com/v1/models` with your key — pick a model from that
52 list, not from this table.
53 Tsubasa implements only `GET /v1/models` and Chat Completions. Its two public
54 model ids share a 32,768-token context, smaller than the 128,000 tokens Codewhale assumes
55 for an unknown model, so set it on the route after saving:
56 `codewhale config set providers.tsubasa.context_window 32768`.
57
58 Cheaper Inference is an OpenAI-compatible gateway with one key for models from
59 several labs. Model ids are bare, such as `gpt-5.4-mini` or `claude-sonnet-5`.
60 Pricing varies by model and route; consult the provider’s current catalog.
61 The authority is `GET https://api.cheaperinference.com/v1/models` with your key.
62 Docs: <https://cheaperinference.com/docs>.
63
64 OpenCode Zen and OpenCode Go are first-class provider routes, configured like
65 any other provider below; they are not part of this table. In `/provider`,
66 type to filter the list (letters not bound to a row action); `Ctrl+T` probes the
67 selected row's `/models` and records reachability only (a 2xx is not
68 model-ready).
69
70 Sources to keep in sync:
71
72 - `crates/config/assets/provider_descriptors.json` - built-in and compatible-host
73 labels, defaults, aliases, key-env lists, config/secret slots, and credential
74 guidance. `crates/config/build.rs` generates immutable typed views and the
75 existing constant projections from this one data owner.
76 - `crates/config/src/provider_kind.rs` and `src/lib.rs` - legacy identity serde,
77 config schema, environment precedence and Rust route/auth behavior.
78 - `crates/tui/src/config.rs` - captures and verifies exact configured provider
79 identities and keeps provider-specific credential and route policy in Rust.
80 - `crates/config/assets/catalog_corrections.json` `reviewed` - intrinsic facts,
81 scoped selector aliases, completion references and pure transport metadata.
82 The existing seed renderer embeds this reviewed supplement in Models.dev.
83 - `crates/agent/src/lib.rs` - compatibility projection of those shared selector
84 rows for `codewhale model list` and `codewhale model resolve`.
85 - `config.example.toml` and `docs/CONFIGURATION.md` - user-facing config
86 examples and environment variable reference.
87 - `scripts/check-provider-registry.py` - drift check for canonical provider
88 IDs, live TUI provider IDs, TOML table names, static registry rows, and
89 documented defaults.
90
91 ## Captured Provider Identity
92
93 Presentation names, labels, aliases, and historical wire tags come from
94 `provider_descriptors.json`. `ProviderKind` remains the Rust-owned intrinsic
95 credential, protocol, and region distinction. A configured route captures both
96 its kind and exact table key; a custom table named `openai` stays custom and
97 does not acquire OpenAI credentials or protocol rules from its name. Case is
98 significant for custom keys.
99
100 DeepSeek China has three preserved representations: `deepseek_c_n` in released
101 TUI serde wrappers, `deepseek-cn` as its captured route ID, and `deepseek_cn` as
102 its config leaf. Its intrinsic kind is DeepSeek, while its table and endpoint
103 stay distinct.
104
105 An old custom route without an additive provider ID can resume only with a
106 verified root `base_url` migration into the active `providers.custom` table.
107 That private receipt is bound to the parsed table generation. A table-only,
108 profile-only, conflicting, or later replaced table cannot establish the
109 missing identity; an explicit empty ID is refused. Writes reuse the existing
110 locked config mutation and undo comparison, obtain the fresh migration receipt
111 under that lock, and verify the captured table before changing a leaf.
112
113 In-flight requests retain their captured identity, endpoint, and credential
114 generation. Health lookup cannot treat an opaque credential reference as ready
115 merely because an earlier request used the same authentication class.
116
117 ## Provider Selection
118
119 With no saved model, no `default_text_model`, and no `CODEWHALE_MODEL` or
120 provider-specific model variable, a fresh install runs `deepseek-flash` on the
121 DeepSeek provider. Precedence is the active provider's configured default, then
122 its catalog default — so an explicit `default_text_model` is not silently
123 overridden by whichever model the shipped catalog lists first.
124
125 Refresh model catalogs without installing a new Codewhale release:
126
127 ```sh
128 codewhale models --update
129 codewhale models --update --provider openai
130 codewhale models --provider openai --json
131 ```
132
133 `models --update` (also `--refresh`) updates the shared Models.dev metadata
134 and calls the existing `/models` endpoint for each configured provider with
135 its own credentials. `--provider ID` restricts the refresh to that exact
136 provider, including named custom endpoints. It makes no inference requests
137 and never changes the saved provider or model. A command-line API key is
138 confined to the active provider; other routes are reported as skipped for
139 that invocation.
140
141 Plain `models` lists the active provider's saved catalog without provider requests.
142 For ChatGPT, it validates the selected registration locally before reading
143 that account's saved roster. Successful refreshes are saved under Codewhale's
144 catalog directory and used by the model/provider pickers. Cache files are
145 scoped to provider identity and endpoint; a failed refresh preserves prior
146 rows. Text output reports source, last successful fetch time (Unix seconds),
147 and freshness. `--update --json` adds per-source receipts and aggregate counts;
148 partial failures return a nonzero exit code after writing those receipts.
149 Ordinary `models --json` keeps its model-array format. Bundled/configured
150 fallbacks are not proof that an account can use every listed model.
151
152 ### Sign in with ChatGPT
153
154 The `openai-codex` provider ID now uses OpenAI's official open-source
155 [Sign in with ChatGPT flow](https://developers.openai.com/siwc/token-sharing-open-source).
156 Sign in and refresh the selected account's catalog:
157
158 ```sh
159 codewhale auth chatgpt
160 codewhale models --update --provider openai-codex
161 codewhale --provider openai-codex
162 ```
163
164 Sign-in opens your system browser, uses a loopback callback and PKCE, validates
165 the returned identity, and stores the issued registration and renewable tokens
166 in Codewhale's protected credential storage. Permission to use your ChatGPT
167 plan is separate from identity sign-in. A declined or missing plan grant stops
168 inference; choose another provider explicitly if you want another billing path.
169 `codewhale auth chatgpt-revoke` signs out of the Codewhale-owned session.
170 This experimental implementation retains one active registration. A picker
171 for multiple saved ChatGPT accounts is not implemented yet.
172
173 Model discovery uses your granted bearer token at
174 `GET https://api.openai.com/v1/models`; inference uses the public
175 `POST https://api.openai.com/v1/responses` endpoint. The model picker keeps
176 OpenAI's display labels and order and lists only entries with `visibility: list`.
177 The catalog contains no credentials and is bound to the verified issuer,
178 issued client ID, and account subject. Another account or workspace never
179 inherits those model choices. Missing, invalid, or older-than-one-day rosters
180 are reported explicitly and provide no account entitlement evidence. Run the
181 refresh command after signing in. Neither model discovery nor ordinary listing
182 starts an inference request. Imported Codex CLI tokens and legacy process-token
183 variables cannot authorize this official plan route.
184
185 Eligible requests consume your ChatGPT plan or credits. ChatGPT Plus's
186 five-hour allowance is shared across apps; each app receives no separate
187 allowance. The documented five-hour limit does not apply to Pro. App-specific
188 limits can also apply. Review limits and access in
189 [ChatGPT usage settings](https://chatgpt.com/settings/usage). Codewhale does
190 not silently switch to an API key or another provider when a limit is reached.
191
192 This is an OpenAI preview for open-source/local apps. It does not grant access
193 to ChatGPT conversation history. Requests stream with `store: false` and
194 `stream: true`, while Codewhale retains its own session history and tools.
195 OpenAI-hosted image generation, file search, Code Interpreter, native computer
196 use, hosted MCP/connectors, and Responses `tool_search` are unavailable on this
197 route; Codewhale's own tools use supported function/custom tool calls.
198 The ChatGPT preview requests one function call at a time.
199 See [preview limitations](https://developers.openai.com/siwc/token-sharing-open-source/preview-limitations).
200 Paid or remotely hosted applications require the
201 [commercial partner interest process](https://openai.com/form/sign-in-with-chatgpt-interest/);
202 this local integration is not commercial approval.
203
204 The canonical provider IDs are the entries of `ProviderKind::ALL`
205 (`crates/config/src/provider_kind.rs`), in that order:
206
207 `deepseek`, `nvidia-nim`, `openai`, `atlascloud`, `wanjie-ark`, `volcengine`,
208 `openrouter`, `orcarouter`, `xiaomi-mimo`, `novita`, `fireworks`, `siliconflow`, `arcee`,
209 `siliconflow-CN`, `moonshot`, `sglang`, `vllm`, `ollama`, `ollama-cloud`, `huggingface`,
210 `together`, `qianfan`, `openai-codex`, `anthropic`, `openmodel`, `zai`,
211 `stepfun`, `minimax`, `deepinfra`, `sakana`, `longcat`, `opencode-go`,
212 `opencode-zen`, `meta`, `xai`, `mistral`, `telecomjs`, `modelstudio-token-plan`, `modelscope`,
213 `google`, `edenai`, `zenmux`, `csdn`, `concentrate`, `codewhale`, and `custom`.
214
215 `deepseek-anthropic` is *not* on this list — it is a wire dialect of
216 `deepseek`, reached with `wire = "anthropic"`, not a separate route to select.
217
218 Use any of these surfaces to select a provider:
219
220 - CLI: `codewhale --provider <id>`
221 - TUI: `/provider <id>` or the provider picker
222 - Env: `CODEWHALE_PROVIDER=<id>`; `DEEPSEEK_PROVIDER=<id>` is the legacy alias
223 - Config: `provider = "<id>"`
224
225 `deepseek-cn`, `deepseek_china`, `deepseekcn`, and `deepseek-china` are accepted
226 as legacy aliases for `deepseek`. They do not select a different official host;
227 DeepSeek uses the same official API host worldwide.
228
229 `deepseek_anthropic`, `deepseek-claude`, and `deepseek_claude` select
230 `deepseek-anthropic`, the opt-in DeepSeek route that speaks the Anthropic
231 Messages API at `https://api.deepseek.com/anthropic`. It keeps the normal
232 DeepSeek API key path but uses `x-api-key` plus `anthropic-version: 2023-06-01`
233 instead of Bearer auth. If the key already lives in official DeepSeek Harness
234 (`dsh`) at `$DSH_HOME/.credentials.yaml`, grant read-only access with
235 `codewhale auth external-consent --provider deepseek --mode read-only`.
236 Codewhale never writes that file and only reads `DEEPSEEK_API_KEY`.
237
238 `huggingface`, `hugging-face`, `hugging_face`, and `hf` all select the
239 Hugging Face Inference Providers route. This is the OpenAI-compatible router
240 path for chat/inference, not Hub browsing, model-card inspection, uploads, or
241 artifact export.
242
243 `telecomjs`, `telecom-js`, `telecom_js`, `telecomjs-cn`, and `tokenhub` all
244 select the TelecomJS TokenHub route. Its authenticated `/models` catalog is
245 key-scoped and remains isolated from every other provider's live snapshot.
246
247 Fresh shared config writes to `~/.codewhale/config.toml`. Existing
248 `~/.deepseek/config.toml` files are still read for compatibility.
249
250 ### Legacy Antigravity tombstone
251
252 Antigravity is not a Codewhale provider and cannot be selected or run. Existing
253 legacy Antigravity provider state is recognized only as a non-runnable migration
254 tombstone. Run `codewhale auth clear --provider antigravity` to forget only
255 Codewhale-owned legacy configuration and consent metadata. This does not sign
256 out of, revoke, read, or otherwise alter any official Google or Antigravity
257 session. For Gemini, select the supported `google` provider and supply
258 `GEMINI_API_KEY`.
259
260 ### Wire Protocol Compatibility
261
262 Provider selection is explicit. A model string prefix such as
263 `deepseek-ai/...`, `deepseek/...`, `qwen/...`, or `arcee-ai/...` is a
264 provider-owned wire ID or catalog namespace hint under the selected provider.
265 It is not a provider switch and must not be treated as proof that the route is
266 DeepSeek, OpenRouter, or any other provider.
267
268 Set the route with `provider = "<id>"`, `CODEWHALE_PROVIDER=<id>`, or
269 `codewhale --provider <id>`. Set the request model with `CODEWHALE_MODEL`, a
270 provider-specific model env var, top-level `default_text_model`, or
271 `[providers.<table>].model`. Set the endpoint with `CODEWHALE_BASE_URL`, a
272 provider-specific base URL env var, or `[providers.<table>].base_url`. Set auth
273 with `codewhale auth set --provider <id>`, `[providers.<table>].api_key`, or
274 the listed provider env vars.
275
276 | Provider ID | TOML table | Wire protocol | Auth env vars |
277 | --- | --- | --- | --- |
278 | `deepseek` | `[providers.deepseek]` | Model-aware: Responses default (`deepseek-flash`); Chat Completions (`deepseek-v4-pro`) | `DEEPSEEK_API_KEY` |
279 | `deepseek-anthropic` | `[providers.deepseek_anthropic]` | Anthropic Messages | `DEEPSEEK_API_KEY` |
280 | `nvidia-nim` | `[providers.nvidia_nim]` | OpenAI Chat Completions | `NVIDIA_API_KEY`, `NVIDIA_NIM_API_KEY` |
281 | `openai` | `[providers.openai]` | OpenAI Chat Completions | `OPENAI_API_KEY` |
282 | `atlascloud` | `[providers.atlascloud]` | OpenAI Chat Completions | `ATLASCLOUD_API_KEY` |
283 | `wanjie-ark` | `[providers.wanjie_ark]` | OpenAI Chat Completions | `WANJIE_ARK_API_KEY`, `WANJIE_API_KEY`, `WANJIE_MAAS_API_KEY` |
284 | `volcengine` | `[providers.volcengine]` | OpenAI Chat Completions | `VOLCENGINE_API_KEY`, `VOLCENGINE_ARK_API_KEY`, `ARK_API_KEY` |
285 | `openrouter` | `[providers.openrouter]` | OpenAI Chat Completions | `OPENROUTER_API_KEY` |
286 | `xiaomi-mimo` | `[providers.xiaomi_mimo]` | OpenAI Chat Completions | `XIAOMI_MIMO_TOKEN_PLAN_API_KEY`, `MIMO_TOKEN_PLAN_API_KEY`, `XIAOMI_MIMO_API_KEY`, `XIAOMI_API_KEY`, `MIMO_API_KEY` |
287 | `novita` | `[providers.novita]` | OpenAI Chat Completions | `NOVITA_API_KEY` |
288 | `fireworks` | `[providers.fireworks]` | OpenAI Chat Completions | `FIREWORKS_API_KEY` |
289 | `siliconflow` | `[providers.siliconflow]` | OpenAI Chat Completions | `SILICONFLOW_API_KEY` |
290 | `arcee` | `[providers.arcee]` | OpenAI Chat Completions | `ARCEE_API_KEY` |
291 | `siliconflow-CN` | `[providers.siliconflow_cn]` | OpenAI Chat Completions | `SILICONFLOW_API_KEY` |
292 | `moonshot` | `[providers.moonshot]` | OpenAI Chat Completions | `MOONSHOT_API_KEY`, `KIMI_API_KEY` |
293 | `sglang` | `[providers.sglang]` | OpenAI Chat Completions | `SGLANG_API_KEY` |
294 | `vllm` | `[providers.vllm]` | OpenAI Chat Completions | `VLLM_API_KEY` |
295 | `ollama` | `[providers.ollama]` | Local OpenAI-compatible Chat Completions | `OLLAMA_API_KEY` (optional; only for an authenticated local route) |
296 | `ollama-cloud` | `[providers.ollama_cloud]` | Hosted OpenAI-compatible Chat Completions | `OLLAMA_CLOUD_API_KEY`, `OLLAMA_API_KEY` |
297 | `huggingface` | `[providers.huggingface]` | OpenAI Chat Completions | `HUGGINGFACE_API_KEY`, `HF_TOKEN` |
298 | `modelscope` | `[providers.modelscope]` | OpenAI Chat Completions | `MODELSCOPE_API_KEY` |
299 | `together` | `[providers.together]` | OpenAI Chat Completions | `TOGETHER_API_KEY` |
300 | `qianfan` | `[providers.qianfan]` | OpenAI Chat Completions | `QIANFAN_API_KEY`, `BAIDU_QIANFAN_API_KEY` |
301 | `openai-codex` | `[providers.openai_codex]` | OpenAI Responses | Official Sign in with ChatGPT (`codewhale auth chatgpt`) with a validated Codewhale-owned plan grant |
302 | `anthropic` | `[providers.anthropic]` | Anthropic Messages | `ANTHROPIC_API_KEY` |
303 | `openmodel` | `[providers.openmodel]` | Anthropic Messages | `OPENMODEL_API_KEY` |
304 | `zai` | `[providers.zai]` | OpenAI Chat Completions | `ZAI_API_KEY`, `Z_AI_API_KEY` |
305 | `stepfun` | `[providers.stepfun]` | OpenAI Chat Completions | `STEPFUN_API_KEY`, `STEP_API_KEY` |
306 | `minimax` | `[providers.minimax]` | OpenAI Chat Completions | `MINIMAX_API_KEY` |
307 | `deepinfra` | `[providers.deepinfra]` | OpenAI Chat Completions | `DEEPINFRA_API_KEY`, `DEEPINFRA_TOKEN` |
308 | `sakana` | `[providers.sakana]` | OpenAI Chat Completions | `FUGU_API_KEY`, `SAKANA_API_KEY` |
309 | `longcat` | `[providers.longcat]` | OpenAI Chat Completions | `LONGCAT_API_KEY` |
310 | `opencode-go` | `[providers.opencode_go]` | OpenAI Chat Completions | `OPENCODE_GO_API_KEY` |
311 | `opencode-zen` | `[providers.opencode_zen]` | Model-aware: OpenAI Responses, Anthropic Messages, or OpenAI Chat Completions | `OPENCODE_ZEN_API_KEY`, `OPENCODE_API_KEY` |
312 | `meta` | `[providers.meta]` | OpenAI Chat Completions | `META_MODEL_API_KEY`, `MODEL_API_KEY` |
313 | `telecomjs` | `[providers.telecomjs]` | OpenAI Chat Completions | `TELECOMJS_API_KEY` |
314 | `xai` | `[providers.xai]` | OpenAI Chat Completions | `XAI_API_KEY` |
315 | `mistral` | `[providers.mistral]` | OpenAI Chat Completions | `MISTRAL_API_KEY` |
316 | `google` | `[providers.google]` | OpenAI Chat Completions (official Gemini OpenAI-compat route; captures and replays thought signatures on tool calls) | `GOOGLE_API_KEY`, `GEMINI_API_KEY` |
317 | `edenai` | `[providers.edenai]` | OpenAI Chat Completions | `EDENAI_API_KEY` |
318 | `zenmux` | `[providers.zenmux]` | OpenAI Chat Completions | `ZENMUX_API_KEY` |
319 | `csdn` | `[providers.csdn]` | OpenAI Chat Completions | `CSDN_API_KEY` |
320 | `concentrate` | `[providers.concentrate]` | OpenAI Responses (`/v1/responses`) | `CONCENTRATE_API_KEY` |
321 | `codewhale` | `[providers.codewhale]` | Model-aware: OpenAI Chat Completions (`/v1/chat/completions`) or Anthropic Messages (`/v1/messages`), chosen per model by the account catalog | `CODEWHALE_API_KEY` |
322 | `modelstudio-token-plan` | `[providers.modelstudio_token_plan]` | OpenAI Chat Completions | `MODELSTUDIO_API_KEY`, `DASHSCOPE_API_KEY` |
323 | `modelstudio-token-plan-anthropic` | `[providers.modelstudio_token_plan_anthropic]` | Anthropic Messages | `MODELSTUDIO_API_KEY`, `DASHSCOPE_API_KEY` |
324 | `modelstudio-coding-plan` | `[providers.modelstudio_coding_plan]` | OpenAI Chat Completions | `MODELSTUDIO_API_KEY`, `DASHSCOPE_API_KEY` |
325 | `modelstudio-coding-plan-anthropic` | `[providers.modelstudio_coding_plan_anthropic]` | Anthropic Messages | `MODELSTUDIO_API_KEY`, `DASHSCOPE_API_KEY` |
326
327 Default base URLs and models for each route are listed in the shipped provider
328 table below. The wire protocol values above are derived from
329 `crates/config/src/provider.rs`: `ChatCompletions` is the default,
330 `openai-codex` overrides to `Responses`; `deepseek-anthropic`, `anthropic`, and
331 `openmodel` override to `AnthropicMessages`; `opencode-zen` resolves the
332 protocol from the selected model's curated offering; and `deepseek` is
333 model-aware — the shipped default `deepseek-flash` (and legacy
334 `deepseek-v4-flash`) rides the Responses endpoint while `deepseek-v4-pro`
335 stays on Chat Completions.
336
337 ## Auth And Env Rules
338
339 For hosted providers, `codewhale auth set --provider <id>` saves an API key for
340 that provider. API-key environment variables are fallback inputs after saved
341 config and keyring credentials; an explicit process-level `--api-key` still
342 wins for that launch.
343
344 For base URL and model selection, prefer:
345
346 - `CODEWHALE_BASE_URL` / `CODEWHALE_MODEL` for the active provider.
347 - Provider-specific base URL/model env vars when listed below.
348 - `DEEPSEEK_BASE_URL`, `DEEPSEEK_MODEL`, and `DEEPSEEK_DEFAULT_TEXT_MODEL` as
349 legacy aliases.
350
351 Non-local `http://` base URLs are rejected unless
352 `DEEPSEEK_ALLOW_INSECURE_HTTP=1` is set. Loopback HTTP URLs are allowed for
353 self-hosted runtimes.
354
355 ## Custom DeepSeek-Compatible Endpoints
356
357 Most custom DeepSeek-compatible deployments can use an existing provider ID.
358 Do not create `[providers.deepseek_custom]`; the provider table names are fixed.
359 Instead, choose the closest shipped route and override its endpoint/model:
360
361 - DeepSeek-compatible hosted API: keep `provider = "deepseek"` and set
362 `[providers.deepseek].base_url` plus `[providers.deepseek].model`, or launch
363 with `DEEPSEEK_BASE_URL` and `DEEPSEEK_MODEL`.
364 - Generic OpenAI-compatible gateway: use `provider = "openai"` with
365 `[providers.openai].base_url` plus `[providers.openai].model`, or launch with
366 `OPENAI_BASE_URL` and `OPENAI_MODEL`.
367 - Multiple named OpenAI-compatible gateways, or local routes you want to pin
368 from an AgentProfile, can use a custom table such as
369 `[providers.lm-studio] kind = "openai-compatible"` and select it with
370 `provider = "lm-studio"` or a profile `provider = "lm-studio"`.
371 - Local OpenAI-compatible runtimes: use `provider = "vllm"`, `"sglang"`, or
372 `"ollama"` with the matching provider-specific base URL/model values.
373
374 Example user config for a DeepSeek-compatible host:
375
376 ```toml
377 provider = "deepseek"
378
379 [providers.deepseek]
380 api_key = "YOUR_API_KEY"
381 base_url = "https://your-provider.example/v1"
382 model = "deepseek-ai/DeepSeek-V4-Pro"
383 ```
384
385 Example user config for a generic gateway:
386
387 ```toml
388 provider = "openai"
389
390 [providers.openai]
391 api_key = "YOUR_GATEWAY_API_KEY"
392 base_url = "https://gateway.example/v1"
393 model = "your-deepseek-compatible-model"
394 ```
395
396 Alibaba Cloud Model Studio (Bailian / DashScope) is a first-class provider as
397 of v0.9.4 with two plan profiles: Token Plan (Personal / Team) and Coding Plan.
398 Both plans expose an OpenAI-compatible Chat Completions endpoint and an
399 Anthropic-compatible Messages endpoint.
400
401 **Token Plan** (Personal and Team share the same AP-Southeast endpoint):
402
403 ```toml
404 provider = "modelstudio-token-plan"
405
406 [providers.modelstudio_token_plan]
407 api_key = "YOUR_MODELSTUDIO_API_KEY"
408 # base_url defaults to https://token-plan.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1
409 model = "qwen3.8-max" # or qwen3.8-max-preview | qwen3.7-plus | qwen3.7-max |
410 # qwen3.6-flash | deepseek-v4-pro | deepseek-v4-flash-0731 |
411 # glm-5.2
412 ```
413
414 **Coding Plan** (separate international endpoint):
415
416 ```toml
417 provider = "modelstudio-coding-plan"
418
419 [providers.modelstudio_coding_plan]
420 api_key = "YOUR_MODELSTUDIO_API_KEY"
421 # base_url defaults to https://coding-intl.dashscope.aliyuncs.com/v1
422 model = "qwen3.8-max"
423 ```
424
425 **Anthropic-compatible dialect** — both plans also expose a native Anthropic
426 Messages path. Select it with the `-anthropic` provider suffix:
427
428 ```toml
429 provider = "modelstudio-token-plan-anthropic"
430
431 [providers.modelstudio_token_plan_anthropic]
432 api_key = "YOUR_MODELSTUDIO_API_KEY"
433 # base_url defaults to https://token-plan.ap-southeast-1.maas.aliyuncs.com/apps/anthropic
434 model = "qwen3.8-max"
435 ```
436
437 Create or copy a Model Studio API key from the
438 [Bailian console](https://bailian.console.aliyun.com/). The API key is shared
439 across all four provider IDs above; only the base URL and wire protocol differ.
440
441 **Thinking / reasoning.** Reasoning surfaces in the TUI's Thinking view on both
442 dialects, per Model Studio's
443 [deep-thinking docs](https://www.alibabacloud.com/help/en/model-studio/deep-thinking).
444
445 On the OpenAI-compatible routes the top-level controls are **route- and
446 model-specific**, and Codewhale fails closed: they are sent only when the
447 configured `base_url` is an official Alibaba Chat Completions host
448 (`*.maas.aliyuncs.com/compatible-mode/v1`, including workspace-scoped hosts, or
449 `coding-intl.dashscope.aliyuncs.com/v1`). A custom `base_url` on the same
450 provider ID gets `thinking`, `enable_thinking`, `preserve_thinking`, and
451 `reasoning_effort` stripped, so an arbitrary OpenAI-compatible gateway is never
452 handed Alibaba's dialect. On a verified host:
453
454 - **Hybrid models** (`qwen3.7-*`, `qwen3.6-*`, `deepseek-v4*`, `glm-*`,
455 `kimi-k2.6*`) get `enable_thinking`: `false` for `off`, `true` otherwise.
456 - **Thinking-only models** — `qwen3.8-max` (catalogued `thinking: always_on`),
457 `qwen3.8-max-preview` (effort/budget options, no toggle), and
458 `kimi-k2.7-code` — get **no** enable/disable switch at all. Sending one is at
459 best ignored.
460 - `preserve_thinking` is sent for the models documented to accept it
461 (`qwen3.7-max`/`-plus`, `qwen3.6-max-preview`/`-plus`/`-flash`, `kimi-k2.6*`,
462 `kimi-k2.7-code`), so the next turn keeps the assistant's trace.
463 - `reasoning_effort` is sent only for the two families with a documented ladder
464 — `deepseek-v4*` and `glm-5`/`5.1`/`5.2` — mapped to `high` or `max`.
465
466 Reasoning streams back as `delta.reasoning_content`. It is replayed to the
467 provider on later turns only for the `preserve_thinking` models above and the
468 thinking-only models; `deepseek-v3.1`, `deepseek-v3.2`, and `glm-*` history
469 stays stripped pending live confirmation that DashScope accepts
470 `reasoning_content` in input messages. (`deepseek-v4*` replays regardless — the
471 DeepSeek thinking-mode contract requires it on every provider.)
472
473 On the Anthropic-compatible routes, thinking uses the documented
474 `{"type":"enabled","budget_tokens":N}` / `{"type":"disabled"}` shapes from the
475 [Anthropic-compatible Messages API](https://www.alibabacloud.com/help/en/model-studio/anthropic-api-messages),
476 with `budget_tokens` derived from the effort level.
477
478 DeepSeek (`deepseek-v4-pro`, `deepseek-v4-flash-0731`) and GLM (`glm-5.2`)
479 models served by Model Studio are provider-scoped and do not collide with the
480 first-party DeepSeek or Zhipu/Z.ai routes. Model Studio publishes no `glm-5.3`
481 entry, so Codewhale does not offer one on this route.
482 Pay-as-you-go workspace-id templating is not yet in the built-in provider; use
483 a custom provider entry for that plan until a follow-up adds it.
484
485 Private gateways with broken or intercepted certificates should use
486 `SSL_CERT_FILE` with a trusted CA bundle. The legacy
487 `insecure_skip_tls_verify = true` key is still parsed so `codewhale doctor` can
488 report stale configs, but provider clients reject it instead of skipping TLS
489 certificate verification.
490
491 Keep `provider`, `api_key`, and `base_url` in user config or process
492 environment. Project-local config overlays intentionally cannot set those keys,
493 so a repository cannot silently redirect prompts or credentials to another
494 endpoint.
495
496 ## Local Models (DS4, Ollama, vLLM, SGLang)
497
498 Self-hosted OpenAI-compatible runtimes are first-class routes and are keyless
499 by default — set an API key only when your server requires one. Start your
500 runner, then point Codewhale at it with `--provider` / `/provider` or a config
501 table.
502
503 | Runner | Default base URL | Default model | Base URL override |
504 | --- | --- | --- | --- |
505 | `ollama` | `http://localhost:11434/v1` | live tag from `GET /v1/models` (pre-refresh: `unknown`) | `OLLAMA_BASE_URL` |
506 | `vllm` | `http://localhost:8000/v1` | `deepseek-ai/DeepSeek-V4-Pro` | `VLLM_BASE_URL` |
507 | `sglang` | `http://localhost:30000/v1` | `deepseek-ai/DeepSeek-V4-Pro` | `SGLANG_BASE_URL` |
508
509 ### DS4 (DwarfStar)
510
511 [DS4](https://github.com/antirez/ds4/tree/84cc882352757baf628a1776badf7cc54d584e28)
512 serves DeepSeek V4 Flash and Pro locally
513 through an OpenAI-compatible API. Start DS4, then open Codewhale's prefilled,
514 keyless setup form:
515
516 ```bash
517 ./ds4-server --ctx 100000 --kv-disk-dir /tmp/ds4-kv --kv-disk-space-mb 8192
518 codewhale
519 # In Codewhale: /setup provider ds4
520 ```
521
522 Review the prefilled route and press Enter to save it. The preset budgets a
523 100,000-token context to match that starter command and defaults to the Flash
524 compatibility alias. Check the local route explicitly with
525 `codewhale doctor --probe-local`.
526
527 DS4 loads the actual GGUF when the server starts. Its `deepseek-v4-flash` and
528 `deepseek-v4-pro` API ids are compatibility aliases; changing `/model` does
529 not swap the resident model. To run Pro, download the supported Pro weights
530 and start `ds4-server -m <pro.gguf> ...` as described by DS4. Update
531 `context_window` whenever the server's `--ctx` value changes.
532
533 The equivalent config is:
534
535 ```toml
536 provider = "ds4"
537
538 [providers.ds4]
539 kind = "openai-compatible"
540 base_url = "http://127.0.0.1:8000/v1"
541 model = "deepseek-v4-flash"
542 auth_mode = "none"
543 context_window = 100000
544 ```
545
546 Codewhale reuses its existing OpenAI-compatible transport and DeepSeek
547 reasoning/tool-call shaping for DS4. It does not invent an API key, confuse an
548 API alias with the loaded GGUF, or silently switch to a hosted DeepSeek route.
549 The pinned DS4 [agent-client contract](https://github.com/antirez/ds4/blob/84cc882352757baf628a1776badf7cc54d584e28/README.md#agent-client-usage)
550 documents Chat Completions at `/v1`, DeepSeek thinking replay, streamed usage,
551 `max_tokens`, and no strict-tool mode; Codewhale follows those exact route
552 facts instead of inheriting unsupported capabilities from a generic gateway.
553 The primary sources for model-facing behavior are DeepSeek's official
554 [thinking-mode](https://api-docs.deepseek.com/guides/thinking_mode),
555 [tool-call](https://api-docs.deepseek.com/guides/tool_calls), and
556 [Chat Completion](https://api-docs.deepseek.com/api/create-chat-completion)
557 contracts. The pinned
558 [DeepSeek Harness adapter](https://github.com/deepseek-ai/deepseek-harness/blob/47f943859bef60e4160492346772ded9b24f765a/packages/llm/llm-deepseek/README.md)
559 is only a secondary implementation cross-check; it is not the API contract.
560
561 ### Ollama
562
563 ```bash
564 ollama serve # if not already running
565 ollama pull <model> # e.g. deepseek-v4-flash, or any tag you prefer
566 codewhale --provider ollama --model <model>
567 ```
568
569 Provider-hinted model names are sent as-is, so `--model qwen3:8b` works with
570 any tag Ollama has pulled.
571
572 ### Ollama Cloud
573
574 Ollama Cloud is a separate hosted provider. It uses the authenticated
575 OpenAI-compatible `/v1/chat/completions` route and defaults to `gpt-oss:120b`:
576
577 ```toml
578 provider = "ollama-cloud"
579
580 [providers.ollama_cloud]
581 base_url = "https://ollama.com/v1"
582 model = "gpt-oss:120b"
583 ```
584
585 Create a key in [Ollama account settings](https://ollama.com/settings/keys),
586 then run `codewhale auth set --provider ollama-cloud`. For ambient auth,
587 `OLLAMA_CLOUD_API_KEY` wins over Ollama's official `OLLAMA_API_KEY`.
588 `OLLAMA_CLOUD_BASE_URL` and `OLLAMA_CLOUD_MODEL` override the Cloud defaults;
589 arbitrary provider-owned model IDs pass through unchanged. Local `ollama`
590 remains a separate, keyless-by-default provider.
591
592 Compatibility is read-only and in memory: a released config that selected
593 `provider = "ollama"` with the exact normalized
594 `[providers.ollama] base_url = "https://ollama.com/v1"` tuple is treated as
595 `ollama-cloud` at runtime. Only that exact tuple may fall back to the legacy
596 `ollama` secret slot. Codewhale does not rewrite the config, copy or delete a
597 secret, migrate neighboring paths, or make an explicit `ollama-cloud` route
598 consume the legacy slot.
599
600 ### vLLM
601
602 ```bash
603 vllm serve <model> --port 8000
604 # or: python -m vllm.entrypoints.openai.api_server --model <model> --port 8000
605 codewhale --provider vllm --model <model>
606 ```
607
608 vLLM's OpenAI-compatible server listens on port 8000 by default, matching
609 Codewhale's `VLLM_BASE_URL`.
610
611 ### SGLang
612
613 ```bash
614 python -m sglang.launch_server --model-path <model> --port 30000
615 codewhale --provider sglang --model <model>
616 ```
617
618 SGLang's default port 30000 matches Codewhale's `SGLANG_BASE_URL`.
619
620 ### Pinning a local route in config
621
622 ```toml
623 provider = "ollama" # or "vllm" / "sglang"
624
625 [providers.ollama]
626 model = "qwen3:8b" # default is deepseek-v4-flash
627 # base_url defaults to http://localhost:11434/v1
628 ```
629
630 Local models that print tool-call JSON without the wire markers: see
631 [When a Local Model Prints Tool JSON](#when-a-local-model-prints-tool-json).
632
633 ## Credential Links
634
635 Provider setup surfaces use the same typed credential metadata as onboarding,
636 `/provider`, `/links`, setup receipts, and doctor output. A missing URL is
637 intentional: local, OAuth-only, and user-defined routes show their supported
638 configuration path instead of guessing a vendor page.
639
640 | Provider ID | Credential or console link |
641 | --- | --- |
642 | `deepseek`, `deepseek-anthropic` | [DeepSeek API keys](https://platform.deepseek.com/api_keys) |
643 | `nvidia-nim` | [NVIDIA NIM API keys](https://build.nvidia.com/settings/api-keys) |
644 | `openai` | [OpenAI API keys](https://platform.openai.com/api-keys) |
645 | `atlascloud` | [Atlas Cloud API keys](https://atlascloud.ai/docs/en/api-keys) |
646 | `wanjie-ark` | [Wanjie MaaS APIKEY docs](https://docs.wanjiedata.com/maas/maas-openapi-v1.html) |
647 | `volcengine` | [Volcengine Ark API keys](https://console.volcengine.com/ark/apiKey) |
648 | `openrouter` | [OpenRouter keys](https://openrouter.ai/settings/keys) |
649 | `xiaomi-mimo` | [Xiaomi MiMo Token Plan](https://platform.xiaomimimo.com/token-plan) |
650 | `novita` | [Novita key management](https://novita.ai/en/settings/key-management) |
651 | `fireworks` | [Fireworks API keys](https://fireworks.ai/api-keys) |
652 | `siliconflow` | [SiliconFlow global API keys](https://cloud.siliconflow.com/account/ak) |
653 | `siliconflow-CN` | [SiliconFlow China API keys](https://cloud.siliconflow.cn/account/ak) |
654 | `arcee` | [Arcee API key guide](https://docs.arcee.ai/other/create-your-first-api-key) |
655 | `moonshot` | [Kimi API platform keys](https://platform.kimi.ai/console/api-keys) or [Kimi Code membership console](https://www.kimi.com/code/console) |
656 | `zai` | [Z.ai model API](https://z.ai/model-api) |
657 | `stepfun` | [StepFun Open Platform](https://platform.stepfun.ai/) |
658 | `minimax`, `minimax-anthropic` | [MiniMax interface keys](https://platform.minimax.io/user-center/basic-information/interface-key) |
659 | `huggingface` | [Hugging Face tokens](https://huggingface.co/settings/tokens) |
660 | `modelscope` | [ModelScope API Keys](https://modelscope.cn/my/settings/token) |
661 | `deepinfra` | [DeepInfra API keys](https://deepinfra.com/dash/api_keys) |
662 | `together` | [Together API keys](https://api.together.ai/settings/api-keys) |
663 | `qianfan` | [Baidu Cloud access keys](https://console.bce.baidu.com/iam/#/iam/accesslist) |
664 | `anthropic` | [Anthropic API keys](https://console.anthropic.com/settings/keys) |
665 | `openmodel` | [OpenModel console](https://console.openmodel.ai/) ([authentication guide](https://docs.openmodel.ai/en/docs/getting-started/authentication)) |
666 | `openai-codex` | Official Sign in with ChatGPT via `codewhale auth chatgpt` (eligible ChatGPT plan or credits, Codewhale-owned tokens). The `openai` API-key route has separate billing. Legacy imported Codex credentials do not authorize this route. |
667 | `sglang`, `vllm` | Local OpenAI-compatible endpoints are keyless by default; configure a key only when the server requires one. |
668 | `ollama` | Local Ollama is keyless by default; configure a key only when the local server requires one. |
669 | `ollama-cloud` | Create an [Ollama API key](https://ollama.com/settings/keys), save it with `codewhale auth set --provider ollama-cloud`, or set `OLLAMA_CLOUD_API_KEY` / `OLLAMA_API_KEY` in that precedence order. |
670 | `sakana` | [Sakana AI API keys](https://console.sakana.ai/api-keys) ([get started](https://console.sakana.ai/get-started)) |
671 | `longcat` | [Meituan LongCat platform](https://longcat.chat/platform) |
672 | `opencode-go` | [OpenCode Go](https://opencode.ai/docs/go/) |
673 | `opencode-zen` | [OpenCode Zen](https://opencode.ai/docs/zen/) |
674 | `meta` | [Meta Model API](https://developer.meta.com/ai/) |
675 | `telecomjs` | [TelecomJS TokenHub](https://aigw.telecomjs.com/) |
676 | `xai` | [xAI Console](https://console.x.ai/) for an API key, Codewhale-owned device login, or explicitly consented read-only Grok CLI credentials. |
677 | `mistral` | [Mistral Console (la Plateforme)](https://console.mistral.ai/api-keys) |
678 | `google` | [Google AI Studio](https://aistudio.google.com/apikey) — Codewhale uses the official Gemini OpenAI-compatible endpoint and never reads Google OAuth files. |
679 | `edenai` | [Eden AI API keys](https://app.edenai.run/settings/api-keys) |
680 | `zenmux` | [ZenMux API keys](https://zenmux.ai/platform/pay-as-you-go) |
681 | `csdn` | [CSDN 星图 console](https://ai.csdn.net/workbench/api-key) — choose the Coding Plan key type for the `glm_for_coding` plan route; a general key bills metered. Docs: [Coding Plan](https://ai.csdn.net/coding-plan). |
682 | `concentrate` | [Concentrate dashboard](https://concentrate.ai/) → API Keys → Create API Key (a Universal API key); docs: [API introduction](https://concentrate.ai/docs/api-reference/introduction). BYOK only — the key stays in the local secret store. |
683 | `modelstudio-token-plan`, `modelstudio-token-plan-anthropic`, `modelstudio-coding-plan`, `modelstudio-coding-plan-anthropic` | [Alibaba Cloud Model Studio (Bailian console)](https://bailian.console.aliyun.com/) — create or copy a Model Studio API key. |
684 | `codewhale` | [Codewhale account settings](https://app.codewhale.net/settings?section=api) — create an API key with the `models:infer` scope, or run `codewhale account api-keys create --name <name> --use`. |
685 | `custom` | Set the named provider's `base_url` and `api_key_env` or `api_key`; no canonical vendor credential page exists. |
686
687 For Kimi, the official [quickstart](https://platform.kimi.ai/docs/overview)
688 directs users to sign in, open **API Keys**, create and copy a key, and keep it
689 secret. Codewhale links straight to that console and accepts the copied key.
690 It never probes or impersonates `kimi_cli`/`kimi_code_cli`; first-class Kimi
691 OAuth remains blocked on a vendor-registered Codewhale identity.
692
693 ### Subscription sign-in accounts (ChatGPT, xAI)
694
695 Codewhale keeps one sign-in of its own per subscription provider. It is
696 separate from the browser, the ChatGPT or Grok apps, and the Codex or Grok
697 CLIs, so those can be signed in to a different account. A successful login
698 replaces the previous owned grant; there is no picker for multiple saved
699 ChatGPT accounts.
700
701 - **See which account is signed in.** `codewhale auth status --provider
702 openai-codex` (or `--provider xai`) and `codewhale auth list` show an account
703 label when the ID token contains an email. ChatGPT plan and workspace
704 details appear only when those claims are present; a workspace label uses
705 the first characters of its account ID. The label is display metadata, not
706 proof of plan entitlement. A valid sign-in may have no email label. Status
707 reads local credential state without contacting the issuer: an expired
708 access token with a stored refresh token remains structurally usable, while
709 missing or invalid storage, or an expired token without a refresh token,
710 is reported unusable. The provider picker shows the label of the sign-in
711 the route would use (for xAI this includes a consented Grok CLI import).
712 Successful login reports the account when known and whether it replaced
713 an earlier sign-in, in the TUI's selected language. No token is printed.
714 - **Reauthorize ChatGPT.** Run `/auth chatgpt` inside Codewhale to update the
715 running session, or `codewhale auth chatgpt` from a shell and restart open
716 Codewhale sessions. Ordinary login reuses the selected account/workspace's
717 issued client ID and checks that the returned account is the same. The
718 browser opens with `prompt=login`; `CODEWHALE_CHATGPT_OAUTH_NO_PROMPT=1`
719 omits that parameter if the issuer refuses it.
720 - **Replace the ChatGPT account or workspace.** From a shell, run
721 `CODEWHALE_CHATGPT_NEW_ACCOUNT=1 codewhale auth chatgpt`, then refresh its
722 model catalog and restart open Codewhale sessions. If the browser holds the
723 wrong account, sign out there or open the printed URL in a private window.
724 Only a fully validated new grant replaces the active owned sign-in.
725 - **Switch xAI accounts.** Run `/auth xai-device` inside Codewhale, or
726 `codewhale auth xai-device` from a shell and restart open sessions. The
727 device page approves with whichever account the browser holds; sign out
728 there or use a private window to approve another account. The login names
729 the account it replaced, or says it is the same account as before.
730 - **Usage limits.** ChatGPT plan requests can fail with
731 `subscription_sharing_usage_limit_exceeded` or
732 `subscription_sharing_usage_unavailable`. These errors, and subscription
733 quota errors from other providers, name the account that made the request
734 when its label is known and offer recovery guidance. Codewhale does not
735 retry these terminal errors or silently choose another billing route. An
736 xAI route set to OAuth that fell back to an API key gets no sign-in
737 guidance, since no sign-in was used.
738 - **ChatGPT credential authority.** The official ChatGPT plan route uses
739 only its selected, verified Codewhale-owned grant.
740 `OPENAI_CODEX_ACCESS_TOKEN`, `CODEX_ACCESS_TOKEN`, and consented Codex CLI
741 tokens cannot override or authorize this route.
742
743 ### External CLI credential consent
744
745 Credential files owned by another CLI are disabled by default. Without an
746 explicit grant, provider discovery, setup, routing, `auth status`, and doctor
747 do not stat, read, refresh, contact an identity provider for, or rewrite Codex,
748 Grok, Kimi, or future external credential files.
749
750 Codewhale currently supports exact-path, provider-scoped **read-only** grants
751 for the Codex CLI and Grok CLI. Codex grants remain available for legacy
752 credential inspection only; they cannot authorize the official ChatGPT plan
753 route. Use `codewhale auth chatgpt` for that route.
754
755 ```bash
756 # Legacy Codex credential inspection; does not authorize ChatGPT plan requests.
757 codex login
758 codewhale auth external-consent --provider openai-codex --mode read-only
759
760 grok login
761 codewhale auth external-consent --provider xai --mode read-only
762
763 codewhale auth status --provider openai-codex
764 codewhale auth external-revoke --provider openai-codex
765 ```
766
767 Pass `--path /absolute/path/to/auth.json` when the external CLI uses a custom
768 location. Consent persists the provider, external owner, exact absolute path,
769 and consent schema version. Later environment-variable changes do not redirect
770 that authority to a different file. Read-only grants never refresh, contact an
771 identity/discovery service, or rewrite the external file. Supported routes such
772 as xAI may use the external token after explicit selection. The official ChatGPT
773 plan route ignores imported Codex tokens and requires its own verified grant.
774 An expired external token fails with login guidance. Doctor reports structural
775 consent/config state without
776 opening credential files and is always non-mutating.
777
778 `managed` is reserved for a future provider-specific preservation adapter.
779 v0.9.1 rejects it before file or network I/O because no reviewed adapter can
780 yet preserve every unknown external schema field safely. Codewhale-started xAI
781 device login instead atomically activates a Codewhale-owned generation named
782 `$CODEWHALE_HOME/credentials/xai-auth-<generation>.json`, stores only that
783 validated basename in config, and revokes any Grok-file grant. Superseded
784 generations are cleaned only after the new config pointer commits.
785 Kimi remains API-key-only; external consent for Kimi is rejected.
786
787 The official DeepSeek Harness (`dsh`) is a third read-only credential owner:
788 `codewhale auth external-consent --provider deepseek --mode read-only` grants
789 exact-path read access to `DEEPSEEK_API_KEY` in `$DSH_HOME/.credentials.yaml`
790 (or `~/.dsh/.credentials.yaml`), which Codewhale never writes, refreshes, or
791 loads into the process environment. This is separate from the DSH *harness*
792 integration (`codewhale integrations dsh …`, see
793 [INTEGRATIONS_DSH.md](INTEGRATIONS_DSH.md)), which never touches credentials
794 in either direction: it pins Codewhale's route identity into a `--patch`
795 overlay and lets DSH resolve its own keys.
796
797 ## Shipped Providers
798
799 | Provider ID | TOML table | Auth env | Base URL env and default | Default or static models | Notes |
800 | --- | --- | --- | --- | --- | --- |
801 | `deepseek` | `[providers.deepseek]` | `DEEPSEEK_API_KEY` | `CODEWHALE_BASE_URL` / `DEEPSEEK_BASE_URL`; default `https://api.deepseek.com/beta` | `deepseek-flash` (shipped default; V4.1 Flash, unversioned id), `deepseek-v4-pro`, `deepseek-v4-flash`, experimental `deepseek-v4-flash-vision-exp`; vision aliases `flash-vision`, `deepseek-v4flashvisionexp`; compatibility aliases `deepseek-chat`, `deepseek-reasoner` | First-class default. The live Pro backend is labeled `DeepSeek-V4-Pro-0813`; the callable API ID remains `deepseek-v4-pro`. Beta URL enables strict tool mode, chat prefix completion, and FIM completion. The documented V4 routes can use provider-native web search through a separate bounded Responses request; compatible custom endpoints do not inherit that capability. Set `https://api.deepseek.com` or `/v1` explicitly to opt out of beta-only features. The shipped default `deepseek-flash` speaks the Responses API (DeepSeek's documented path for Codex-style integration, since the 2026-07-31 Flash production update); the Chat-only controls (the `thinking` toggle and strict-tool `/beta` routing) apply to Chat Completions routes, and `deepseek-v4-pro` stays on Chat Completions until its announced Responses rollout. Reasoning effort follows DeepSeek's documented requested-to-actual mapping: `minimal`/`low` land on `low`, `medium`/`xhigh` on `high`, and `max`/`ultra` on `max`; `off` disables thinking (`thinking: {"type":"disabled"}` on Chat, `reasoning.effort: "none"` on Responses). The experimental vision ID was observed in the authenticated `/models` roster on 2026-08-21 and is advertised as image-input capable on the direct Chat Completions route only. Its limits, reasoning, and tool-call flags provisionally inherit Flash; pricing remains unknown, and no funded image round trip was made during this release work. |
802 | `deepseek-anthropic` | `[providers.deepseek_anthropic]` | `DEEPSEEK_API_KEY` | `DEEPSEEK_ANTHROPIC_BASE_URL`; default `https://api.deepseek.com/anthropic` | `deepseek-v4-pro`, `deepseek-v4-flash`; compatibility aliases `deepseek-chat`, `deepseek-reasoner` | Opt-in DeepSeek route for the Anthropic Messages wire protocol. Uses `/v1/messages`, `x-api-key`, and `anthropic-version: 2023-06-01`. Keep `provider = "deepseek"` for the default Chat Completions path. |
803 | `nvidia-nim` | `[providers.nvidia_nim]` | `NVIDIA_API_KEY`, `NVIDIA_NIM_API_KEY` | `NVIDIA_NIM_BASE_URL`, `NIM_BASE_URL`, `NVIDIA_BASE_URL`; default `https://integrate.api.nvidia.com/v1` | `deepseek-ai/deepseek-v4-pro`, `deepseek-ai/deepseek-v4-flash` | Hosted DeepSeek V4 through NVIDIA NIM. `NVIDIA_NIM_MODEL` is accepted by the TUI config path. |
804 | `openai` | `[providers.openai]` | `OPENAI_API_KEY` | `OPENAI_BASE_URL`; default `https://api.openai.com/v1` | `gpt-5.6` (default), `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna` | Generic OpenAI-compatible route whose built-in endpoint and fallback catalog are native to OpenAI. The [GPT-5.6 family](https://developers.openai.com/api/docs/models/gpt-5.6-sol) uses OpenAI's documented 1.05M context, 128K max output, and reasoning levels. Custom gateways remain free to select an explicit gateway-owned model. `OPENAI_MODEL` is accepted. |
805 | `atlascloud` | `[providers.atlascloud]` | `ATLASCLOUD_API_KEY` | `ATLASCLOUD_BASE_URL`; default `https://api.atlascloud.ai/v1` | Default `deepseek-ai/deepseek-v4-flash`; explicit `vendor/model-id` values pass through when AtlasCloud is selected | OpenAI-compatible hosted route. `ATLASCLOUD_MODEL` is accepted by the TUI config path, the static `ModelRegistry` keeps DeepSeek V4 fallback rows, and provider-hinted CLI model IDs are sent to AtlasCloud exactly as requested. Use Atlas Cloud's own catalog or Coding Plan page for the current provider-owned model list and pricing. |
806 | `wanjie-ark` | `[providers.wanjie_ark]` | `WANJIE_ARK_API_KEY`, `WANJIE_API_KEY`, `WANJIE_MAAS_API_KEY` | `WANJIE_ARK_BASE_URL`, `WANJIE_BASE_URL`, `WANJIE_MAAS_BASE_URL`; default `https://maas-openapi.wanjiedata.com/api/v1` | `deepseek-reasoner` | OpenAI-compatible hosted route. `WANJIE_ARK_MODEL`, `WANJIE_MODEL`, and `WANJIE_MAAS_MODEL` are accepted. |
807 | `volcengine` | `[providers.volcengine]` | `VOLCENGINE_API_KEY`, `VOLCENGINE_ARK_API_KEY`, `ARK_API_KEY` | `VOLCENGINE_BASE_URL`, `VOLCENGINE_ARK_BASE_URL`, `ARK_BASE_URL`; default `https://ark.cn-beijing.volces.com/api/coding/v3` | `DeepSeek-V4-Pro`, `DeepSeek-V4-Flash` | Volcengine/Volcano Engine Ark OpenAI-compatible coding endpoint. `VOLCENGINE_MODEL` and `VOLCENGINE_ARK_MODEL` are accepted. |
808 | `openrouter` | `[providers.openrouter]` | `OPENROUTER_API_KEY` | `OPENROUTER_BASE_URL`; default `https://openrouter.ai/api/v1` | `deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash`; recent large IDs include `arcee-ai/trinity-large-thinking`, `minimax/minimax-m3`, `xiaomi/mimo-v2.5-pro`, `qwen/qwen3.6-flash`, `qwen/qwen3.6-35b-a3b`, `qwen/qwen3.6-max-preview`, `qwen/qwen3.6-27b`, `qwen/qwen3.6-plus`, `google/gemma-4-31b-it`, `z-ai/glm-5.1`, `z-ai/glm-5.2`, `moonshotai/kimi-k2.7-code`, `moonshotai/kimi-k2.6` | Additive open-model routing layer. It does not replace DeepSeek; it lets users route supported model IDs through OpenRouter when they choose it. |
809 | `orcarouter` | `[providers.orcarouter]` | `ORCAROUTER_API_KEY` | `ORCAROUTER_BASE_URL`; default `https://api.orcarouter.ai/v1` | `deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash`; router alias `orcarouter/auto`; recent large IDs mirror the OpenRouter namespaced catalog | [OrcaRouter](https://www.orcarouter.ai) OpenAI-compatible aggregation gateway. Shares the namespaced `vendor/model` wire-model format and DeepSeek model set with OpenRouter, so the OpenRouter base-URL and model-normalization rules apply. `ORCAROUTER_MODEL` is accepted. Provider aliases: `orcarouter`, `orca_router`, `orca`. |
810 | `xiaomi-mimo` | `[providers.xiaomi_mimo]` | `XIAOMI_MIMO_TOKEN_PLAN_API_KEY`, `MIMO_TOKEN_PLAN_API_KEY`, `XIAOMI_MIMO_API_KEY`, `XIAOMI_API_KEY`, `MIMO_API_KEY` | `XIAOMI_MIMO_BASE_URL`, `MIMO_BASE_URL`, `XIAOMI_MIMO_MODE`, `MIMO_MODE`; default `https://token-plan-sgp.xiaomimimo.com/v1` | Chat: `mimo-v2.5-pro`, `mimo-v2.5-pro-ultraspeed`, `mimo-v2.5`; speech/TTS: `mimo-v2.5-tts`, `mimo-v2.5-tts-voicedesign`, `mimo-v2.5-tts-voiceclone`, `mimo-v2-tts` | Xiaomi MiMo OpenAI-compatible chat completions route. `mimo-v2.5-pro` and `mimo-v2.5` can use the documented provider-native web-search plugin; ultraspeed, speech/TTS, and custom-compatible routes do not inherit it. Token Plan keys (`tp-...`) use `api-key` auth and the token-plan endpoint by default; pay-as-you-go mode uses standard API keys (`sk-...`) and `https://api.xiaomimimo.com/v1`. It sends `max_completion_tokens` and uses MiMo's `thinking` field for reasoning control. Token Plan cost/usage is credit/quota based; Codewhale shows it as unknown until Xiaomi exposes a reliable balance API. `codewhale speech` / `tts` uses the TTS models. |
811 | `novita` | `[providers.novita]` | `NOVITA_API_KEY` | `NOVITA_BASE_URL`; default `https://api.novita.ai/openai/v1` | `deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash` | OpenAI-compatible hosted route for DeepSeek model IDs. Use config or `CODEWHALE_MODEL` / `DEEPSEEK_MODEL` for model overrides. |
812 | `fireworks` | `[providers.fireworks]` | `FIREWORKS_API_KEY` | `FIREWORKS_BASE_URL`; default `https://api.fireworks.ai/inference/v1` | `accounts/fireworks/models/deepseek-v4-pro` | OpenAI-compatible hosted route. Use config or `CODEWHALE_MODEL` / `DEEPSEEK_MODEL` for model overrides. |
813 | `siliconflow` | `[providers.siliconflow]` | `SILICONFLOW_API_KEY` | `SILICONFLOW_BASE_URL`; default `https://api.siliconflow.com/v1` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash` | OpenAI-compatible hosted route. Official docs use the `.com` endpoint. `SILICONFLOW_MODEL` is accepted. Reasoning aliases `deepseek-reasoner` and `deepseek-r1` map to Pro; `deepseek-chat` and `deepseek-v3` map to Flash. |
814 | `siliconflow-CN` | `[providers.siliconflow_cn]` | `SILICONFLOW_API_KEY` | `SILICONFLOW_BASE_URL`; default `https://api.siliconflow.cn/v1` | Uses the SiliconFlow model set | China regional SiliconFlow route. Falls back to `[providers.siliconflow]` for api_key / base_url / model when unset. Select it with `provider = "siliconflow-CN"` or `CODEWHALE_PROVIDER=siliconflow-CN`. |
815 | `arcee` | `[providers.arcee]` | `ARCEE_API_KEY` | `ARCEE_BASE_URL`; default `https://api.arcee.ai/api/v1` | `trinity-large-thinking`, `trinity-large-preview` | Arcee AI direct OpenAI-compatible route, tracked as 256K-context BF16 serving. `ARCEE_MODEL` is accepted. OpenRouter's `arcee-ai/trinity-large-thinking` remains the OpenRouter namespaced model ID; direct Arcee uses the bare `trinity-large-thinking` ID. |
816 | `moonshot` | `[providers.moonshot]` | `MOONSHOT_API_KEY`, `KIMI_API_KEY` | `MOONSHOT_BASE_URL`, `KIMI_BASE_URL`; default `https://api.moonshot.ai/v1` | Direct Moonshot: `kimi-k3`, `kimi-k2.7-code`, `kimi-k2.7-code-highspeed`, `kimi-k2.6`; Kimi Code membership: `k3`, `kimi-for-coding`, `kimi-for-coding-highspeed` at `https://api.kimi.com/coding/v1` | Moonshot/Kimi route. Exact direct `kimi-k3` routes use the documented Formula web-search tool/fiber loop; direct `kimi-k2.6` retains the built-in `$web_search` contract, and exact Kimi Code membership routes use their structured `/search` service. Adjacent paths, K2.7 direct models, and cross-product model IDs do not inherit native search. `kimi` and `kimi-k2` aliases select `kimi-k2.7-code`; `MOONSHOT_MODEL`, `KIMI_MODEL_NAME`, and `KIMI_MODEL` are accepted. Kimi thinking streams through `reasoning_content`; Codewhale keeps it in Thinking cells and replays it for thinking/tool-call continuity. For direct K3, use exact `base_url = "https://api.moonshot.ai/v1"` and `model = "kimi-k3"`; it is always-thinking and receives top-level `reasoning_effort = "low" | "high" | "max"` (`off` normalizes to `low`), uses only `max_completion_tokens`, and omits `temperature`/`top_p` per the [K3 quickstart](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart). For Kimi Code K3, use a key from the [Kimi Code console](https://www.kimi.com/code/console), exact `base_url = "https://api.kimi.com/coding/v1"`, and bare `model = "k3"`; `off` becomes enabled `low`, while normal dispatched `auto` selects and sends a concrete Codewhale tier. Only an omitted reasoning setting leaves the provider default in control. That membership route defaults safely to 262,144 context tokens; the [Kimi Code model-tier table](https://www.kimi.com/code/docs/en/kimi-code/models.html) grants Allegretto and higher plans up to 1M, which those plans may express as `context_window = 1048576`. `k3[1m]` is Claude Code-only and Codewhale rejects it. `kimi-for-coding` remains the valid K2.7 membership route, and `kimi-for-coding-highspeed` is its own high-speed roster entry (262,144 context); membership ids are rejected on the direct platform endpoint, and `kimi-k3` stays rejected on the membership endpoint. Billing is decided by the endpoint the route resolves to, judged once against the two exact product endpoints: direct Moonshot (`https://api.moonshot.ai/v1` or the default) bills metered with dollar estimates, the exact Kimi Code membership endpoint bills as Kimi Code quota and never shows dollar estimates, and anything else — a gateway host, a neighboring Kimi-hosted path — reports `cost: unknown` rather than borrowing either product. An imported Kimi Code token with no `base_url` in its table still resolves to the membership endpoint, so it bills as Kimi Code quota and never accrues dollars. A completed turn, parent or sub-agent, is billed from the immutable endpoint receipt its own client was built with, never from a later config re-read: `MOONSHOT_BASE_URL`/`KIMI_BASE_URL` are merged into the *active* provider's table only, and an in-turn provider switch can move the ambient config off the route that actually ran. Legacy `auth_mode = "kimi_oauth"` fails to API-key guidance without probing Kimi CLI files. Codewhale does not impersonate `kimi_cli` or `kimi_code_cli`. **China-region keys:** contributor field evidence (@vFONGv, PR #5229, verified on Windows 10) reports that a China-region Moonshot key must be paired with `base_url = "https://api.moonshot.cn/v1"`; left on the default international host (`https://api.moonshot.ai/v1`) it fails authentication. We have no China-region key to verify this ourselves, so it is recorded as a user report rather than a tested route. Note also that editing `base_url` alone does not take effect until `codewhale auth set` is re-run for that provider. |
817 | `google` | `[providers.google]` | `GOOGLE_API_KEY`, `GEMINI_API_KEY` | `GOOGLE_BASE_URL`, `GEMINI_BASE_URL`; default `https://generativelanguage.googleapis.com/v1beta/openai/` | `gemini-3.1-pro-preview` (default); `/model` also lists `gemini-3-pro-preview`, `gemini-3.7-flash`, `gemini-3.6-flash`, `gemini-3.5-flash`, `gemini-3.5-flash-lite`, `gemini-2.5-pro`, `gemini-2.5-flash` | Google Gemini as its own backend on the official OpenAI-compatible Chat Completions route. Thinking models capture `extra_content.google.thought_signature` on tool calls and replay it with the assistant tool-call messages; replaying a tool call whose signature was not captured fails closed with an actionable error instead of letting the tool loop break. `gemini-2.5-flash-lite` ships thinking off and degrades with a warning instead. Reasoning effort maps onto the documented `google.thinking_config.thinking_level` (`low`/`high`). The dialect binds to the exact official base URL: a `google` row pointed at another gateway gets plain OpenAI semantics and no signature requirements. Codewhale never reads Google OAuth files; only an AI Studio API key is used. Not live-tested against the real endpoint in this environment. |
818 | `zai` | `[providers.zai]` | `ZAI_API_KEY`, `Z_AI_API_KEY`, `ZHIPU_API_KEY`, `GLM_API_KEY` | `ZAI_BASE_URL`, `Z_AI_BASE_URL`; default `https://api.z.ai/api/coding/paas/v4`; general APIs `https://api.z.ai/api/paas/v4` and `https://open.bigmodel.cn/api/paas/v4` | `GLM-5.3` default; `/model` also lists `GLM-5.3-Flash`, `GLM-5.2`, `GLM-5.1`, and `GLM-5-Turbo` | Z.AI GLM Coding Plan route. All three first-party Chat routes (both api.z.ai products and BigModel's general platform endpoint) share one request dialect, so reasoning controls — the GLM-5.2 thinking toggle, tiered effort, and the forced-thinking GLM-5.3 rewrite that sends `off` as `enabled` + `reasoning_effort: "low"` — apply on BigModel too; neighboring paths such as `/preview` stay fail-closed. The two general API products expose structured provider-native web search (`search-prime` globally, `search_std` in China); Coding Plan and compatible custom endpoints do not inherit it. `GLM-5.3` is the default and a first-class picker row (`model = "GLM-5.3"` or `ZAI_MODEL=GLM-5.3`); `GLM-5.3-Flash` is the 1M multimodal fast sibling (`model = "GLM-5.3-Flash"`). An explicit `GLM-5.2` selection keeps its own id. Limits and reasoning options for 5.3 are inherited from `GLM-5.2` until Z.ai publishes distinct 5.3 metadata; 5.3 carries no price. Flash ships the published $0.15/$0.50 list. A live call can still 429 with entitlement code 1311 on accounts that are not provisioned for 5.3. |
819 | `stepfun` | `[providers.stepfun]` | `STEPFUN_API_KEY`, `STEP_API_KEY` | `STEPFUN_BASE_URL`, `STEP_BASE_URL`; default `https://api.stepfun.ai/v1`; Coding Plan endpoint `https://api.stepfun.ai/step_plan/v1` | `step-3.7-flash` | StepFun / StepFlash direct OpenAI-compatible route. `/provider` setup asks which billing route the key belongs to — pay-as-you-go or Step Plan — validates the key against the chosen endpoint, and writes the answer to `[providers.stepfun].base_url` only. A base URL that is neither recognized route is left alone and the question is skipped. You can also set `[providers.stepfun].base_url` or `STEP_BASE_URL` to the Coding Plan URL by hand. Offline accounting labels recognized routes as `stepfun-payg` or `stepfun-plan` without persisting the raw endpoint, and only the standard PAYG route receives token pricing. `STEPFUN_MODEL` and `STEP_MODEL` are accepted. |
820 | `minimax` | `[providers.minimax]` | `MINIMAX_API_KEY` | `MINIMAX_BASE_URL`; default `https://api.minimax.io/v1`; China `https://api.minimaxi.com/v1` | `MiniMax-M3`, `MiniMax-M2.7`, `MiniMax-M2.7-highspeed`, `MiniMax-M2.5`, `MiniMax-M2.5-highspeed`, `MiniMax-M2.1`, `MiniMax-M2.1-highspeed`, `MiniMax-M2` | MiniMax direct OpenAI-compatible route. Codewhale sends `reasoning_split = true` so MiniMax thinking arrives separately from answer text. Both MiniMax dialects sell pay-as-you-go and Token Plan over the same endpoints and the same key, so billing is classified from the credential *product*, never from the endpoint or from a default. `mode = "token-plan"` in `[providers.minimax]`/`[providers.minimax_anthropic]`, or a Token Plan key shaped `sk-cp…`, bills as MiniMax Token Plan quota with no dollar estimates; an explicit pay-as-you-go mode (`pay-as-you-go`/`payg`/`metered`) wins over key shape. The key's product prefix is only visible when the key is in config, bound by `api_key_env`, or exported as `MINIMAX_API_KEY` on an official endpoint — a key saved through `codewhale auth set` (secret store / OS keyring) is deliberately not read to classify billing. With no explicit mode and no visible product marker the route reports `cost: unknown` rather than assuming pay-as-you-go, so a Token Plan account is never charged invented dollars. Custom/gateway endpoints also fail closed with `cost: unknown`. Official M3 input modalities are text, image, and video; M2.7 is text-only. |
821 | `minimax-anthropic` | `[providers.minimax_anthropic]` | `MINIMAX_API_KEY` | `MINIMAX_ANTHROPIC_BASE_URL`; default `https://api.minimax.io/anthropic`; China `https://api.minimaxi.com/anthropic` | `MiniMax-M3`, `MiniMax-M2.7` | MiniMax direct Anthropic-compatible Messages route. Keep the `/anthropic` suffix because Codewhale appends `/v1/messages`; the route uses `x-api-key`. M3 supports adaptive or disabled thinking. M2.7 always keeps thinking enabled. |
822 | `sglang` | `[providers.sglang]` | Optional `SGLANG_API_KEY` | `SGLANG_BASE_URL`; default `http://localhost:30000/v1` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash` | Self-hosted OpenAI-compatible route. Localhost deployments commonly omit auth. `SGLANG_MODEL` is accepted. |
823 | `vllm` | `[providers.vllm]` | Optional `VLLM_API_KEY` | `VLLM_BASE_URL`; default `http://localhost:8000/v1` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash` | Self-hosted vLLM OpenAI-compatible route. Localhost deployments commonly omit auth. `VLLM_MODEL` is accepted. |
824 | `ollama` | `[providers.ollama]` | Local optional `OLLAMA_API_KEY` | `OLLAMA_BASE_URL`; default `http://localhost:11434/v1` | live tag from the local catalog; pre-refresh placeholder `unknown`; provider-hinted custom tags pass through | Local Ollama is keyless by default. `OLLAMA_MODEL` is accepted. The header must not paint a hosted id the local daemon did not list. |
825 | `ollama-cloud` | `[providers.ollama_cloud]` | `OLLAMA_CLOUD_API_KEY`, then `OLLAMA_API_KEY` | `OLLAMA_CLOUD_BASE_URL`; default `https://ollama.com/v1` | `gpt-oss:120b`; arbitrary provider-owned IDs pass through | Hosted OpenAI-compatible `/v1/chat/completions` route. Save credentials under `ollama-cloud`; the exact released `ollama` + Cloud URL tuple has bounded read-only in-memory compatibility with its legacy table and secret slot. `OLLAMA_CLOUD_MODEL` is accepted. |
826 | `huggingface` | `[providers.huggingface]` | `HUGGINGFACE_API_KEY`, `HF_TOKEN` | `HUGGINGFACE_BASE_URL`, `HF_BASE_URL`; default `https://router.huggingface.co/v1` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash` | Hugging Face Inference Providers OpenAI-compatible router route. Accepted aliases: `huggingface`, `hugging-face`, `hugging_face`, `hf`. Org-prefixed model IDs pass through. `HUGGINGFACE_MODEL` and `HF_MODEL` are accepted. Hub browsing/export are separate future features. |
827 | `modelscope` | `[providers.modelscope]` | `MODELSCOPE_API_KEY` | `MODELSCOPE_BASE_URL`; default `https://api-inference.modelscope.cn/v1` | `Qwen/Qwen3.5-397B-A17B` (default), `Qwen/Qwen3.5-122B-A10B`, `Qwen/Qwen3.5-27B`, `Qwen/Qwen3.5-35B-A3B`, `Qwen/Qwen3.8-27B`, `Qwen/Qwen3.8-Flash-Next`, `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Pro-0813`, `deepseek-ai/DeepSeek-V4.1-Flash`, `ZhipuAI/GLM-4.7-Flash`, `ZhipuAI/GLM-5.2` | ModelScope OpenAI-compatible inference route. Org-prefixed model IDs pass through. `MODELSCOPE_MODEL` is accepted. |
828 | `deepinfra` | `[providers.deepinfra]` | `DEEPINFRA_API_KEY`, `DEEPINFRA_TOKEN` | `DEEPINFRA_BASE_URL`; default `https://api.deepinfra.com/v1/openai` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash` | DeepInfra OpenAI-compatible route. Drop-in replacement for OpenAI SDK. |
829 | `together` | `[providers.together]` | `TOGETHER_API_KEY` | `TOGETHER_BASE_URL`; default `https://api.together.xyz/v1` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash`, `thinkingmachines/inkling` | Together AI OpenAI-compatible route. `TOGETHER_MODEL` is accepted. Model aliases `deepseek-v4-pro` and `deepseek-v4-flash` normalize to Together's org-prefixed IDs; `inkling` and `together-inkling` normalize to Together's published lowercase Inkling wire ID. Inkling uses the exact `none`/`minimal`/`low`/`medium`/`high`/`max` reasoning vocabulary from Thinking Machines' [official model repository](https://huggingface.co/thinkingmachines/Inkling). Together's [launch post](https://www.together.ai/blog/together-ai-brings-thinking-machines-labs-new-model-inkling-on-day-0) currently says Inkling is live with 1M context, while its [model detail page](https://www.together.ai/models/inkling) says coming soon with 256K context and publishes no price. Until Together's active `/models` endpoint and the Models.dev catalog resolve that conflict, Inkling is not seeded into Codewhale's offline picker and no route-specific context or cost is inferred. |
830 | `qianfan` | `[providers.qianfan]` | `QIANFAN_API_KEY`, `BAIDU_QIANFAN_API_KEY` | `QIANFAN_BASE_URL`, `BAIDU_QIANFAN_BASE_URL`; default `https://api.baiduqianfan.ai/v1` | `ernie-4.0-turbo-8k`; provider-scoped custom Qianfan service/model IDs pass through | Baidu Qianfan OpenAI-compatible route. Requests use Bearer auth and Chat Completions payloads. `QIANFAN_MODEL` and `BAIDU_QIANFAN_MODEL` are accepted; aliases `baidu-qianfan`, `baidu_qianfan`, and `baidu` resolve to this provider. Tool/function calling is model-scoped in Qianfan docs, so Codewhale preserves the selected wire model and leaves live capability proof to follow-up route/capability work. |
831 | `openai-codex` | `[providers.openai_codex]` | Official Sign in with ChatGPT (`codewhale auth chatgpt` / `/provider setup openai-codex`) | Official `https://api.openai.com/v1` | Selected account catalog; configured model IDs remain explicit selections | **Experimental.** Public Responses endpoint (`/v1/responses`) with a validated `chatgpt.tokens.use.direct` grant. Dynamic OSS registration issues a client ID for the selected account/workspace; Codewhale protects and renews its own tokens. The account catalog supplies selectable models; public model lists and legacy Codex tokens do not establish plan permission. No silent billing fallback. See [Sign in with ChatGPT](#sign-in-with-chatgpt) for setup, usage limits, and preview boundaries. |
832 | `anthropic` | `[providers.anthropic]` | `ANTHROPIC_API_KEY` | `ANTHROPIC_BASE_URL`; default `https://api.anthropic.com` | `claude-opus-4-8`, `claude-sonnet-4-6` (default), `claude-haiku-4-5` | Native Anthropic Messages API route (`/v1/messages`, `x-api-key` + `anthropic-version: 2023-06-01`) — not OpenAI-compatible. Prompt caching via `cache_control` breakpoints, adaptive thinking + `output_config.effort`, signed thinking blocks replayed verbatim, cache telemetry normalized per #2961. `ANTHROPIC_MODEL` is accepted. |
833 | `openmodel` | `[providers.openmodel]` | `OPENMODEL_API_KEY` | `OPENMODEL_BASE_URL`; default `https://api.openmodel.ai` | `deepseek-v4-flash`; provider-scoped custom model IDs pass through | OpenModel Anthropic-compatible Messages route. Uses `/v1/messages`, Bearer auth, and `anthropic-version: 2023-06-01`; OpenModel selects DeepSeek, DashScope, Xiaomi, Claude, and other routes by model id. `OPENMODEL_MODEL` is accepted. |
834 | `sakana` | `[providers.sakana]` | `FUGU_API_KEY`, `SAKANA_API_KEY` | `SAKANA_BASE_URL`; default `https://api.sakana.ai/v1` | `fugu` (default), `fugu-ultra-20260615` | Sakana AI Fugu OpenAI-compatible route. Standard Chat Completions wire protocol; streaming supported. `fugu-ultra-20260615` is the heavy/reasoning variant. Env var aliases: `FUGU_API_KEY` (primary), `SAKANA_API_KEY`; provider aliases: `sakana-ai`, `sakana_ai`, `fugu`. |
835 | `longcat` | `[providers.longcat]` | `LONGCAT_API_KEY` | `LONGCAT_BASE_URL`; default `https://api.longcat.chat/openai/v1` | `LongCat-2.0` (default) | Meituan LongCat curated model gateway. OpenAI-compatible Chat Completions wire protocol. Sign up at https://longcat.chat/platform for an API key. Provider aliases: `long-cat`, `meituan-longcat`, `meituan`. |
836 | `opencode-go` | `[providers.opencode_go]` | `OPENCODE_GO_API_KEY` | `OPENCODE_GO_BASE_URL`; default `https://opencode.ai/zen/go/v1` | `deepseek-v4-pro` (default), `grok-4.5`, `glm-5.2`, `glm-5.1`, `kimi-k3`, `kimi-k2.7-code`, `kimi-k2.6`, `deepseek-v4-flash`, `mimo-v2.5`, `mimo-v2.5-pro` | [OpenCode Go](https://opencode.ai/docs/go/) subscription route using OpenAI-compatible Chat Completions. `OPENCODE_GO_MODEL` is accepted. Codewhale uses bare wire IDs; familiar `opencode-go/<model-id>` input aliases normalize to the bare ID. Go models documented only on the Anthropic `/messages` endpoint are deliberately not advertised by this route until Codewhale supports per-model wire selection. Billing surfaces show the Go allowance instead of token-price estimates. |
837 | `opencode-zen` | `[providers.opencode_zen]` | `OPENCODE_ZEN_API_KEY`, fallback `OPENCODE_API_KEY` | `OPENCODE_ZEN_BASE_URL`; default `https://opencode.ai/zen/v1` | `gpt-5.6` (default); current GPT, Claude, Qwen, DeepSeek, MiniMax, GLM, Kimi, Grok, Muse Spark, and free-model IDs | [OpenCode Zen](https://opencode.ai/docs/zen/) model-aware gateway. `OPENCODE_ZEN_MODEL` is accepted, and official `opencode/<model-id>` selectors normalize to bare wire IDs. Each model's wire comes from the curated snapshot, then from the AI SDK package its [Models.dev](https://models.dev) row names: GPT, Grok, and Muse Spark rows use `/responses`; Claude and most Qwen rows use `/messages`; DeepSeek, MiniMax, GLM, Kimi, `qwen3.8-max`, and the free rows use `/chat/completions`. Responses and Chat Completions authenticate with Bearer `Authorization`, while Anthropic Messages uses `x-api-key`; none of these routes use ChatGPT/Codex OAuth guidance or headers. Gemini fails closed because its model-specific Google wire protocol is not implemented; Models.dev rows marked `deprecated` and models no loaded catalog lists also fail closed. See [OpenCode Zen protocol catalog](#opencode-zen-protocol-catalog). |
838 | `meta` | `[providers.meta]` | `META_MODEL_API_KEY`, `MODEL_API_KEY` | `META_MODEL_API_BASE_URL`, `MODEL_API_BASE_URL`; default `https://api.meta.ai/v1` | `muse-spark-1.2` (default) | [Meta Model API](https://developer.meta.com/ai/resources/blog/build-with-muse-spark/) public-preview route using OpenAI-compatible Chat Completions. Muse Spark 1.2 keeps its wire ID, tool support, 1M context, 32K output metadata, and `none` through `xhigh` reasoning effort. `META_MODEL_API_MODEL` and `MODEL_API_MODEL` are accepted. Provider aliases: `meta-ai`, `meta_model_api`, `muse`, `muse-spark`. |
839 | `telecomjs` | `[providers.telecomjs]` | `TELECOMJS_API_KEY` | `TELECOMJS_BASE_URL`; default `https://aigw.telecomjs.com/v1` | `deepseek-v4-pro` conservative fallback; authenticated `/models` rows when a key is configured | TelecomJS TokenHub OpenAI-compatible Chat Completions route. Live catalogs are isolated by provider and key fingerprint, stale rows survive transient refresh failures, and unsupported reasoning request fields are omitted. `TELECOMJS_MODEL` is accepted. Provider aliases: `telecom-js`, `telecom_js`, `telecomjs-cn`, `tokenhub`. |
840 | `mistral` | `[providers.mistral]` | `MISTRAL_API_KEY` | `MISTRAL_BASE_URL`; default `https://api.mistral.ai/v1` | `mistral-code-latest` (default; `codestral-latest` accepted as alias), `mistral-medium-latest` (aliases: `mistral-medium-3-5`), `mistral-small-latest` (aliases: `mistral-small-2603`), `mistral-large-latest` | Mistral AI (la Plateforme) OpenAI-compatible Chat route. On the documented first-party HTTPS `/v1` hosts, Medium and Small send adjustable `reasoning_effort` (`none` or `high` only), parse Mistral's polymorphic thinking/text blocks, and replay stored thinking in that same wire shape. Deprecated native Magistral IDs remain explicit-configuration compatibility routes: they are always-reasoning and never receive the adjustable effort field. Code and Large are non-reasoning. A custom `MISTRAL_BASE_URL` keeps generic Chat semantics unless it is one of the documented first-party hosts. `MISTRAL_MODEL` is accepted. Provider aliases: `mistral-ai`, `mistralai`, `la-plateforme`. |
841 | `edenai` | `[providers.edenai]` | `EDENAI_API_KEY` | `EDENAI_BASE_URL`; default `https://api.edenai.run/v3`; EU `https://api.eu.edenai.run/v3` | `deepseek/deepseek-v4-pro` (default); live `/models` catalog of `provider/model` ids | Eden AI OpenAI-compatible aggregation gateway. Catalog rows remain provider-scoped; generic reasoning controls are omitted because supported fields depend on the selected upstream family. `EDENAI_MODEL` is accepted. The default `deepseek/deepseek-v4-pro` is listed on the global catalog only; on the EU endpoint set `EDENAI_MODEL` (or `model`) to a row from the EU `/models` list, for example `qwen/deepseek-v4-pro`. Provider aliases: `eden-ai`, `eden_ai`. |
842 | `zenmux` | `[providers.zenmux]` | `ZENMUX_API_KEY` | `ZENMUX_BASE_URL`; default `https://zenmux.ai/api/v1` | `deepseek/deepseek-v4.1-flash` (default); live `/models` catalog of `provider/model` ids (keyless-readable) | ZenMux OpenAI-compatible aggregation gateway (~190 models). Catalog rows remain provider-scoped; generic reasoning controls are omitted because supported fields depend on the selected upstream family. `ZENMUX_MODEL` is accepted. Provider aliases: `zen-mux`, `zen_mux`. |
843 | `csdn` | `[providers.csdn]` | `CSDN_API_KEY` | `CSDN_BASE_URL`; default `https://ai.csdn.net/api/model/v1` | `glm_for_coding` (default; the Coding Plan's dedicated model id); other marketplace model ids pass through | CSDN 星图 (Starmap) OpenAI-compatible hosted platform. Coding Plan keys and general marketplace keys share the one endpoint, so billing follows the credential product, never the URL alone: routing `glm_for_coding` — the shipped default — or setting `mode = "coding_plan"`/`"plan"`/`"subscription"` bills as CSDN Coding Plan quota with no dollar estimates; any other model, or an explicit `pay-as-you-go`/`metered` mode, bills metered; an unrecognized mode or an endpoint off `ai.csdn.net/api/model/v1` reports `cost: unknown`. `CSDN_MODEL` is accepted. Provider aliases: `csdn-ai`, `csdn_ai`, `csdn-coding-plan`, `csdn_coding_plan`, `starmap`. |
844 | `concentrate` | `[providers.concentrate]` | `CONCENTRATE_API_KEY` | `CONCENTRATE_BASE_URL`; default `https://api.concentrate.ai/v1` | `deepseek-v4-pro` (default; a plain catalog id lets the gateway pick the upstream provider); `provider/model` ids such as `openai/gpt-5.6-sol` pin one upstream; `concentrate/auto` is the gateway's own router; unauthenticated live `/v1/models` catalog | Concentrate OpenAI Responses-compatible gateway (`POST /v1/responses`, bearer Universal API key). Opt-in and BYOK only: your key, your Concentrate bill, zero Codewhale fee, no managed default. Requests carry only documented fields (the system prompt rides as a `system` input item). Contract: [API introduction](https://concentrate.ai/docs/api-reference/introduction), [request parameters](https://concentrate.ai/docs/api-reference/endpoint/request-parameters), [streaming](https://concentrate.ai/docs/api-reference/endpoint/streaming), [errors](https://concentrate.ai/docs/api-reference/endpoint/errors). See [Concentrate Notes](#concentrate-notes). |
845 | `codewhale` | `[providers.codewhale]` | `CODEWHALE_API_KEY` | `CODEWHALE_API_BASE`; default `https://api.codewhale.net/v1` | `deepseek/deepseek-v4-pro` (default), `anthropic/claude-sonnet-5`, `openai/gpt-5.6` are offline bootstrap rows only; the authenticated `GET /v1/models` listing of the account's connected providers is the catalog authority | Codewhale API: account-backed model access over the provider keys the customer connected to their Codewhale account. One `cwc_key_…` account API key with the `models:infer` scope authenticates every model. Model ids are `provider/model` exactly as the account catalog returns them, and each row states its protocol (`chat-completions` → `/v1/chat/completions`, `anthropic-messages` → `/v1/messages`, `responses` → `/v1/responses`); every protocol uses `Authorization: Bearer`, never `x-api-key`. `CODEWHALE_API_BASE` must be HTTPS except on loopback. Connect provider keys with `codewhale account keys set <provider>`. Provider aliases: `codewhale-api`, `cw-api`, `codewhale-cloud`. |
846 | `xai` | `[providers.xai]` | `XAI_API_KEY`, Codewhale-owned device OAuth, or explicit read-only Grok CLI consent | `XAI_BASE_URL`; default `https://api.x.ai/v1` | `grok-4.6` (default), `grok-4.7`, `grok-4.5`, `grok-4.3`, `grok-build`, `grok-composer-2.5-fast`, `grok-4.20-0309-reasoning`, `grok-4.20-0309-non-reasoning` | xAI/Grok OpenAI-compatible Chat Completions route. Grok 4.6 has a 500K context window, text/image input, function calls, structured output, server-side web search, and `low`/`medium`/`high`/`xhigh` reasoning (default `high`). [Grok 4.7](https://docs.x.ai/docs/models/grok-4.7) has the same 500K window, text/image input, function calls, structured output, reasoning ladder and $2.00 / $0.50 cached / $6.00 rates; xAI documents no web-search support for it yet, so Codewhale does not claim it. Grok reasoning cannot be disabled ([reasoning guide](https://docs.x.ai/docs/guides/reasoning)). Its standard rates double when the prompt reaches 200K tokens; the same 2x long-context rule applies to `grok-4.5` (500K context, $2.00 / $0.30 cached / $6.00) and `grok-4.3` (1M context, $1.25 / $0.20 cached / $2.50) per their [model pages](https://docs.x.ai/docs/models/grok-4.5). There is no documented `latest`/`fast` alias and no published numeric output limit. **API-key** (default): Bearer token from console.x.ai via `XAI_API_KEY` / keyring / `api_key`. **OAuth**: `codewhale auth xai-device` uses SSH-friendly device login and Codewhale-owned storage, which may refresh itself. Existing Grok CLI credentials require `codewhale auth external-consent --provider xai --mode read-only`; the granted external file is never refreshed or rewritten. OAuth may return HTTP 403 on some SuperGrok tiers — keep API-key as the reliable fallback. `XAI_MODEL` is accepted. Provider aliases: `x-ai`, `x_ai`, `grok`. |
847 | `modelstudio-token-plan` | `[providers.modelstudio_token_plan]` | `MODELSTUDIO_API_KEY`, `DASHSCOPE_API_KEY` | `MODELSTUDIO_TOKEN_PLAN_BASE_URL`; default `https://token-plan.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1` | `qwen3.8-max` (default), `qwen3.8-max-preview`, `qwen3.7-plus`, `qwen3.7-max`, `qwen3.6-flash`, `deepseek-v4-pro`, `deepseek-v4-flash-0731`, `glm-5.2` | Alibaba Cloud Model Studio Token Plan OpenAI-compatible Chat Completions route. Token Plan Personal and Team share this endpoint. `qwen3.8-max`, `qwen3.7-plus`, and `qwen3.7-max` can use provider-native web search through the Token Plan Responses Harness; the preview, Coding Plan, and Anthropic routes do not inherit that capability. All listed models are reasoning-capable text/coding models. DeepSeek and GLM entries are provider-scoped and do not collide with first-party routes. `MODELSTUDIO_TOKEN_PLAN_MODEL` is accepted. Provider aliases: `modelstudio-token-plan`, `alibaba-token-plan`, `dashscope-token-plan`. |
848 | `modelstudio-token-plan-anthropic` | `[providers.modelstudio_token_plan_anthropic]` | `MODELSTUDIO_API_KEY`, `DASHSCOPE_API_KEY` | default `https://token-plan.ap-southeast-1.maas.aliyuncs.com/apps/anthropic` | Same model catalog as `modelstudio-token-plan` | Token Plan Anthropic-compatible Messages route (`/apps/anthropic`). Same API key as the OpenAI dialect. Provider aliases: `modelstudio-token-plan-anthropic`, `alibaba-token-plan-anthropic`. |
849 | `modelstudio-coding-plan` | `[providers.modelstudio_coding_plan]` | `MODELSTUDIO_API_KEY`, `DASHSCOPE_API_KEY` | `MODELSTUDIO_CODING_PLAN_BASE_URL`; default `https://coding-intl.dashscope.aliyuncs.com/v1` | `qwen3.8-max` (default); same catalog as Token Plan | Alibaba Cloud Model Studio Coding Plan OpenAI-compatible Chat Completions route. `MODELSTUDIO_CODING_PLAN_MODEL` is accepted. Provider aliases: `modelstudio-coding-plan`, `alibaba-coding-plan`, `dashscope-coding-plan`. |
850 | `modelstudio-coding-plan-anthropic` | `[providers.modelstudio_coding_plan_anthropic]` | `MODELSTUDIO_API_KEY`, `DASHSCOPE_API_KEY` | default `https://coding-intl.dashscope.aliyuncs.com/apps/anthropic` | Same model catalog as `modelstudio-coding-plan` | Coding Plan Anthropic-compatible Messages route (`/apps/anthropic`). Provider aliases: `modelstudio-coding-plan-anthropic`, `alibaba-coding-plan-anthropic`. |
851
852 StepFun's four coding models are available through both its standard API and
853 [Step Plan](https://platform.stepfun.ai/docs/en/step-plan/integrations/reasoning-api).
854 Choose the billing route in `/provider`, then the model in `/model`; existing
855 Step 3.7 selections remain unchanged. [Step 5 Preview](https://platform.stepfun.ai/docs/en/guides/models/step-5-preview)
856 has a 1M context window. Step 5 and Step 3.7 expose low/medium/high reasoning;
857 Step 3.5 Flash 2603 exposes low/high; base Step 3.5 uses provider-default reasoning.
858 [Published API prices](https://platform.stepfun.ai/docs/en/guides/pricing/details)
859 apply only to verified PAYG routes. Step Plan displays subscription allowance.
860 Speech, music and image-generation models use separate interfaces and are not
861 presented as coding models. Provider video capability metadata does not imply
862 that every Codewhale client can attach video.
863
864
865 ### OpenCode Zen protocol catalog
866
867 Zen Responses and Chat Completions requests authenticate with Bearer
868 `Authorization`; Zen Anthropic Messages requests use `x-api-key`. None of these
869 routes add ChatGPT/Codex OAuth headers.
870
871 Each Zen model's wire comes from two sources, and neither is guessed from a
872 model-id family (`qwen3.8-flash` is Messages while `qwen3.8-max` is Chat):
873
874 1. **The curated snapshot** compiled into Codewhale, verified against the
875 [official endpoint table](https://opencode.ai/docs/zen/) and the `opencode`
876 provider in [Models.dev](https://models.dev). It is the offline floor, and
877 it wins when a catalog row names a different wire for the same id.
878 2. **The Models.dev catalog.** Its `opencode` provider is Zen's published
879 catalog: each model row names the AI SDK package OpenCode itself uses, which
880 Codewhale maps exactly — `@ai-sdk/openai` to Responses, `@ai-sdk/anthropic`
881 to Anthropic Messages, and `@ai-sdk/openai-compatible` (the provider
882 default) to Chat Completions. A Zen model released after this build routes
883 once the catalog lists it, without a Codewhale release. The interactive TUI
884 and `codewhale exec` both load the persisted catalog; refresh it with
885 `codewhale models --update`.
886
887 The curated snapshot:
888
889 - Responses: `gpt-6-astra`, `gpt-6-sol`, `gpt-6-luna`, `gpt-5.6-sol`,
890 `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-5.5`, `gpt-5.5-pro`, `gpt-5.4`,
891 `gpt-5.4-pro`, `gpt-5.4-mini`, `gpt-5.4-nano`, `gpt-5.3-codex`,
892 `gpt-5.3-codex-spark`, `gpt-5.2`, `gpt-5.2-codex`, `gpt-5.1`,
893 `gpt-5.1-codex`, `gpt-5.1-codex-max`, `gpt-5.1-codex-mini`, `gpt-5`,
894 `gpt-5-codex`, `gpt-5-nano`, `grok-4.7`, `grok-4.6`, `grok-4.5`,
895 `grok-build-0.1`, `muse-spark-1.3`, `muse-spark-1.3-contributor-free`,
896 `muse-spark-1.2`, `muse-spark-1.2-contributor`,
897 `muse-spark-1.2-contributor-free`.
898 - Anthropic Messages: `claude-fable-5-1`, `claude-fable-5`,
899 `claude-opus-5-5`, `claude-opus-5`, `claude-opus-4-8`, `claude-opus-4-7`,
900 `claude-opus-4-6`, `claude-opus-4-5`, `claude-sonnet-5-5`,
901 `claude-sonnet-5`, `claude-sonnet-4-6`, `claude-sonnet-4-5`,
902 `claude-sonnet-4`, `claude-haiku-4-5`, `qwen3.8-flash`, `qwen3.7-max`,
903 `qwen3.7-plus`, `qwen3.6-plus`, `qwen3.5-plus`.
904 - Chat Completions: `deepseek-v4.1-flash`, `deepseek-v4-pro`,
905 `deepseek-v4-flash`, `deepseek-v4-flash-vision-exp`, `minimax-m3`,
906 `minimax-m2.7`, `minimax-m2.5`, `glm-5.3-flash`, `glm-5.3`, `glm-5.2`,
907 `glm-5.1`, `glm-5`, `kimi-k3`, `kimi-k2.7-code`, `kimi-k2.6`, `kimi-k2.5`,
908 `qwen3.8-max`, `big-pickle`, `space-bunny-free`,
909 `longcat-2.5-preview-free`, `mimo-v2.6-flash-free`, `mimo-v2.5-free`,
910 `ling-3.0-flash-fin-free`, `north-mini-code-free`,
911 `nemotron-3-ultra-free`, `nemotron-3.5-lightning-free`,
912 `deepseek-v4-flash-free`.
913
914 These fail closed locally, with the reason in the error, instead of reaching
915 Zen: Gemini (`@ai-sdk/google`, Google's model-specific protocol, which Codewhale
916 does not speak); a catalog row naming any other package; a catalog row Models.dev
917 marks `deprecated`; and a model neither source lists. A miss never falls back to
918 another Zen wire shape, including when a custom Zen base URL is configured.
919
920 ### Concentrate Notes
921
922 Concentrate (`concentrate`) is an OpenAI **Responses**-compatible gateway:
923 `POST https://api.concentrate.ai/v1/responses` with `Authorization: Bearer
924 <Universal API key>`. It is opt-in and **BYOK only**.
925
926 - **Model ids pass through.** A plain catalog id (`gpt-5.6-sol`,
927 `deepseek-v4-pro`) lets the gateway choose the upstream provider;
928 `provider/model` (`openai/gpt-5.6-sol`) pins the upstream provider inside the
929 gateway; `concentrate/auto` sends the gateway's own `auto` router. Codewhale's
930 bare `--model auto` remains the resolver sentinel (provider default), which is
931 why the gateway router has its own explicit spelling. Once streaming begins
932 the gateway commits to one provider ([auto routing](https://concentrate.ai/docs/api-reference/endpoint/auto-routing)).
933 - **Wire.** Fixed Responses dialect. The request carries only documented
934 fields — `model`, `input`, `stream`, `max_output_tokens`, `tools`,
935 `tool_choice`, `parallel_tool_calls`, `reasoning.effort` — and the system
936 prompt travels as a leading `system` input item because `instructions` is
937 not in the gateway's [parameter reference](https://concentrate.ai/docs/api-reference/endpoint/request-parameters).
938 Streaming follows the typed `response.*` SSE events and ends on
939 `response.completed` without a `[DONE]` sentinel ([streaming](https://concentrate.ai/docs/api-reference/endpoint/streaming)).
940 - **Errors.** `{ "error", "message", "model"?, "retry_after"? }` with 400
941 (bad request / unknown model), 401 (invalid key), 402 (insufficient
942 credits), 424 (upstream provider unavailable), 429 (rate limit) —
943 surfaced verbatim and classified by message ([errors](https://concentrate.ai/docs/api-reference/endpoint/errors)).
944 - **Catalog.** `GET /v1/models` needs no key
945 ([list models](https://concentrate.ai/docs/api-reference/endpoint/list-models)); the live rows
946 are provider-scoped and unclaimed (no pricing or capability claims).
947 - **Commercial boundary.** Concentrate's [Terms of Service](https://concentrate.ai/legal/terms-of-service)
948 forbid reselling, white-labeling, or providing the service on a
949 service-bureau basis without written consent, and its
950 [Acceptable Use Policy](https://concentrate.ai/legal/acceptable-use-policy)
951 forbids sharing or sublicensing keys. Codewhale therefore ships **no**
952 Codewhale-owned Concentrate key, no stored customer key, no default or
953 managed routing, and no markup — activation of any hosted lane is gated on
954 written consent, terms, and billing approval (tracked in the Codewhale ops
955 evidence checklist `concentrate-gateway-20260829/CHECKLIST.md`). Public
956 pricing states "No platform markup on tokens".
957 - **Self-test without a key.** `scripts/concentrate-selftest.sh` boots a
958 local stub that speaks the documented contract and drives the real
959 `codewhale exec` path through it, asserting the URL, bearer header, model
960 passthrough, streaming, and the completed-turn receipt. No network call
961 leaves the machine and no account is required.
962
963 ### Hugging Face Provider vs MCP vs Hub
964
965 Codewhale's `huggingface` provider ID is only the OpenAI-compatible chat
966 inference route through Hugging Face Inference Providers. It is selected with
967 `/provider huggingface`, `CODEWHALE_PROVIDER=huggingface`, or
968 `provider = "huggingface"`.
969
970 Hugging Face MCP is a separate external-tool route. Configure it through the
971 MCP config described in `docs/MCP.md`, preferably using the settings-generated
972 snippet from <https://huggingface.co/settings/mcp>. In the TUI, `/hf mcp status`
973 checks whether the Hugging Face MCP server appears in the resolved MCP config,
974 `/hf mcp setup` prints the settings workflow and a placeholder-only shape, and
975 `/hf concepts` explains the provider/MCP/Hub distinction.
976
977 Hub publishing or repository management remains explicit user action through
978 Hub-native tooling such as `huggingface_hub` or git. The `/hf` helper does not
979 upload to Hugging Face and does not perform direct Hugging Face Hub HTTP search.
980
981 ### Xiaomi MiMo Notes
982
983 `xiaomi-mimo` defaults to `mimo-v2.5-pro` for long-context reasoning and coding
984 work. The chat picker also exposes `mimo-v2.5-pro-ultraspeed` and the latest
985 Omni model `mimo-v2.5`. Xiaomi MiMo TTS is available through
986 `codewhale --provider xiaomi-mimo speech "text" --model tts` (or the `tts`
987 alias). In Act and Operate, the provider-specific `speech` / `tts` tools are
988 available through deferred discovery when the Xiaomi MiMo route is configured.
989
990 `/provider xiaomi-mimo ultraspeed` and `/provider xiaomi-mimo pro-ultraspeed`
991 both select `mimo-v2.5-pro-ultraspeed`. Speech aliases such as `tts`,
992 `voice-design`, and `voice-clone` are separate from normal chat defaults.
993
994 Token Plan keys default to the Singapore endpoint
995 `https://token-plan-sgp.xiaomimimo.com/v1`. If your MiMo account is provisioned
996 for the China region, set `base_url = "https://token-plan-cn.xiaomimimo.com/v1"`
997 explicitly in `[providers.xiaomi_mimo]` or set `mode = "token-plan-cn"`. Europe
998 Token Plan accounts can set
999 `base_url = "https://token-plan-ams.xiaomimimo.com/v1"` or use
1000 `mode = "token-plan-ams"`; `mode = "pay-as-you-go"`
1001 selects the standard API endpoint and standard MiMo key family. Xiaomi Token
1002 Plan docs and console expose credit/quota semantics, but Codewhale does not
1003 currently have a documented balance endpoint to poll, so cost display remains
1004 unknown rather than reusing token-price estimates from another provider.
1005
1006 Voice-design and voice-clone shorthands map to `mimo-v2.5-tts-voicedesign` and
1007 `mimo-v2.5-tts-voiceclone`. Xiaomi's current
1008 [image-understanding guide](https://platform.xiaomimimo.com/docs/en-US/usage-guide/multimodal-understanding/image-understanding)
1009 includes `mimo-v2.5` for image input. Codewhale exposes image analysis through the
1010 separate `[vision_model]` / `image_analyze` path; set that model to
1011 `mimo-v2.5` when using MiMo for vision.
1012
1013 ### OpenRouter-Compatible Base URLs
1014
1015 OpenRouter-compatible gateways should usually stay on the `openrouter`
1016 provider with a provider-scoped `base_url` override instead of moving through
1017 the generic `openai` route. That keeps OpenRouter-style reasoning, streaming,
1018 cache usage, and namespaced wire model parsing attached to the selected route:
1019
1020 ```toml
1021 provider = "openrouter"
1022
1023 [providers.openrouter]
1024 api_key = "sk-..."
1025 base_url = "https://openrouter-compatible.example/v1"
1026 model = "deepseek/deepseek-v4-pro"
1027 ```
1028
1029 Codewhale preserves the `deepseek/` wire-model prefix under the OpenRouter
1030 provider scope; it does not infer a switch to the direct DeepSeek provider from
1031 that model string. Cache fields such as `prompt_cache_hit_tokens`,
1032 `prompt_cache_miss_tokens`, and `prompt_tokens_details.cached_tokens` are
1033 parsed when the upstream gateway sends them. If a key/account type omits those
1034 fields, Codewhale treats them as absent for that response rather than as a
1035 different provider route.
1036
1037 OrcaRouter (`https://api.orcarouter.ai/v1`) is a dedicated named route
1038 ([OrcaRouter](https://www.orcarouter.ai)) that speaks the same OpenAI
1039 Chat Completions wire protocol and serves the same namespaced
1040 `vendor/model` catalog. It does not need an OpenRouter-compatible
1041 `base_url` override: select `provider = "orcarouter"` and its namespaced
1042 wire models (for example `deepseek/deepseek-v4-pro` or its own
1043 `orcarouter/auto` router) pass through verbatim, exactly as they do on the
1044 OpenRouter provider scope.
1045
1046 ### Recent OpenRouter Large Models
1047
1048 OpenRouter completions and static registry rows include the April 2026 onward
1049 large models verified through OpenRouter's model metadata:
1050 `arcee-ai/trinity-large-thinking`, `qwen/qwen3.6-flash`,
1051 `qwen/qwen3.6-35b-a3b`, `qwen/qwen3.6-max-preview`, `qwen/qwen3.6-27b`,
1052 `qwen/qwen3.6-plus`, `minimax/minimax-m3`, `xiaomi/mimo-v2.5-pro`,
1053 `xiaomi/mimo-v2.5`, `moonshotai/kimi-k2.7-code`, `moonshotai/kimi-k2.6`,
1054 `z-ai/glm-5.1`, `z-ai/glm-5.2`, `z-ai/glm-5-turbo`, `tencent/hy3-preview`,
1055 `google/gemma-4-31b-it`, `google/gemma-4-26b-a4b-it`, and
1056 `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free`.
1057 `minimax/minimax-m3` was added from OpenRouter's May 31, 2026 listing as a 1M
1058 context multimodal model for coding, tool use, and long-horizon agentic work.
1059 `GLM-5.3` is now the default direct Z.AI Coding Plan model; `GLM-5.2` /
1060 `z-ai/glm-5.2` remain available (explicit selections keep their own id),
1061 `GLM-5.1` / `z-ai/glm-5.1` remain available as the smaller model,
1062 `GLM-5.3-Flash` / `z-ai/glm-5.3-flash` is the faster/explore sibling of
1063 `GLM-5.3`, and `GLM-5-Turbo` / `z-ai/glm-5-turbo` remains the faster sibling
1064 of `GLM-5.2`.
1065 `GLM-5.3` / `z-ai/glm-5.3` and `GLM-5.3-Flash` / `z-ai/glm-5.3-flash` are
1066 first-class picker ids on the Z.ai and OpenRouter routes (`/model` after
1067 `/provider zai`, or `model = "GLM-5.3-Flash"`).
1068 Limits and reasoning options for 5.3 are inherited from
1069 `GLM-5.2` until Z.ai publishes distinct 5.3 metadata, and 5.3 carries no
1070 price. Flash ships the published $0.15/$0.50 list. A live call can still
1071 429 with entitlement code 1311 on accounts that are not provisioned for 5.3.
1072
1073 ## Static Model Registry
1074
1075 `codewhale model list` and `codewhale model resolve` project the reviewed
1076 `selections` in `crates/config/assets/catalog_corrections.json` through
1077 `crates/agent/src/lib.rs`. There is no independent Rust model roster. These
1078 ordered aliases and flags are compatibility metadata, not account availability
1079 or executable route permission. This differs from live `/models` discovery.
1080 Use `/models` or `codewhale models` to fetch model IDs from the active API
1081 endpoint when the endpoint supports model listing.
1082
1083 | Provider | Static registry entries | Tool calls | Registry reasoning flag |
1084 | --- | --- | --- | --- |
1085 | `deepseek` | `deepseek-v4-pro`, `deepseek-v4-flash`, `deepseek-v4-flash-vision-exp` | yes | yes |
1086 | `nvidia-nim` | `deepseek-ai/deepseek-v4-pro`, `deepseek-ai/deepseek-v4-flash` | yes | yes |
1087 | `openai` | `deepseek-v4-pro`, `deepseek-v4-flash`, `gpt-5.6`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna` | yes | yes |
1088 | `atlascloud` | `deepseek-ai/deepseek-v4-flash`, `deepseek-ai/deepseek-v4-pro` | yes | yes |
1089 | `wanjie-ark` | `deepseek-reasoner` | yes | yes |
1090 | `volcengine` | `DeepSeek-V4-Pro`, `DeepSeek-V4-Flash` | yes | yes |
1091 | `openrouter` | `deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash`, `arcee-ai/trinity-large-thinking`, `minimax/minimax-m3`, `minimax/minimax-m2.7`, `xiaomi/mimo-v2.5-pro`, `xiaomi/mimo-v2.5`, `qwen/qwen3.6-flash`, `qwen/qwen3.6-35b-a3b`, `qwen/qwen3.6-max-preview`, `qwen/qwen3.6-27b`, `qwen/qwen3.6-plus`, `qwen/qwen3.7-max`, `moonshotai/kimi-k2.7-code`, `moonshotai/kimi-k2.6`, `z-ai/glm-5.1`, `z-ai/glm-5.2`, `z-ai/glm-5.3`, `z-ai/glm-5.3-flash`, `z-ai/glm-5-turbo`, `tencent/hy3-preview`, `google/gemma-4-31b-it`, `google/gemma-4-26b-a4b-it`, `nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free`, `nvidia/nemotron-3-ultra-550b-a55b` | yes | yes |
1092 | `orcarouter` | `deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash`, `orcarouter/auto` | yes | yes |
1093 | `xiaomi-mimo` | `mimo-v2.5-pro`, `mimo-v2.5-pro-ultraspeed`, `mimo-v2.5`; speech/TTS IDs are selected through `codewhale speech` / `tts` | yes | yes for chat models; no for speech/TTS models |
1094 | `novita` | `deepseek/deepseek-v4-pro`, `deepseek/deepseek-v4-flash` | yes | yes |
1095 | `fireworks` | `accounts/fireworks/models/deepseek-v4-pro` | yes | yes |
1096 | `siliconflow` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash` | yes | yes |
1097 | `arcee` | `trinity-large-thinking`, `trinity-large-preview`; provider-hinted custom model IDs pass through | yes | yes for `trinity-large-thinking`; no for `trinity-large-preview` |
1098 | `moonshot` | `kimi-k2.7-code`, `kimi-k2.6` | yes | yes |
1099 | `zai` | `GLM-5.3`, `GLM-5.3-Flash`, `GLM-5.2`, `GLM-5.1`, `GLM-5-Turbo`; provider-hinted custom model IDs pass through | yes | yes |
1100 | `stepfun` | `step-3.7-flash` (default), `step-5-preview`, `step-3.5-flash`, `step-3.5-flash-2603` | yes | yes |
1101 | `minimax` | `MiniMax-M3`, `MiniMax-M2.7`, `MiniMax-M2.7-highspeed`, `MiniMax-M2.5`, `MiniMax-M2.5-highspeed`, `MiniMax-M2.1`, `MiniMax-M2.1-highspeed`, `MiniMax-M2` | yes | yes |
1102 | `minimax-anthropic` | `MiniMax-M3`, `MiniMax-M2.7` | yes | yes |
1103 | `sglang` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash` | yes | yes |
1104 | `vllm` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash` | yes | yes |
1105 | `ollama` | live local tag; custom tags pass through when provider hint is `ollama` | yes | no |
1106 | `ollama-cloud` | `gpt-oss:120b`; arbitrary provider-owned model IDs pass through | yes | yes |
1107 | `huggingface` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash` | yes | no |
1108 | `modelscope` | `Qwen/Qwen3.5-397B-A17B`, `Qwen/Qwen3.5-122B-A10B`, `Qwen/Qwen3.5-27B`, `Qwen/Qwen3.5-35B-A3B`, `Qwen/Qwen3.8-27B`, `Qwen/Qwen3.8-Flash-Next`, `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Pro-0813`, `deepseek-ai/DeepSeek-V4.1-Flash`, `ZhipuAI/GLM-4.7-Flash`, `ZhipuAI/GLM-5.2` | yes | no |
1109 | `deepinfra` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash` | yes | yes |
1110 | `together` | `deepseek-ai/DeepSeek-V4-Pro`, `deepseek-ai/DeepSeek-V4-Flash`, `thinkingmachines/inkling` | yes | yes |
1111 | `openai-codex` | `gpt-5.5` | yes | yes |
1112 | `anthropic` | `claude-opus-5`, `claude-opus-4-8`, `claude-sonnet-5`, `claude-sonnet-4-6`, `claude-fable-5`, `claude-haiku-4-5` | yes | yes except `claude-haiku-4-5` |
1113 | `openmodel` | `deepseek-v4-flash`; provider-scoped custom model IDs pass through | yes | model-dependent |
1114 | `sakana` | `fugu`, `fugu-ultra-20260615` | yes | yes for `fugu-ultra-20260615` |
1115 | `longcat` | `LongCat-2.0` | yes | yes |
1116 | `opencode-go` | `deepseek-v4-pro`, `grok-4.5`, `glm-5.2`, `glm-5.1`, `kimi-k3`, `kimi-k2.7-code`, `kimi-k2.6`, `deepseek-v4-flash`, `mimo-v2.5`, `mimo-v2.5-pro` | yes | yes |
1117 | `meta` | `muse-spark-1.2` | yes | yes |
1118 | `xai` | `grok-4.6`, `grok-4.7`, `grok-4.5`, `grok-4.3`, `grok-build`, `grok-composer-2.5-fast`, `grok-4.20-0309-reasoning`, `grok-4.20-0309-non-reasoning` | yes | yes for `grok-4.6`, `grok-4.7`, `grok-4.5`, `grok-4.3`, `grok-build`, and `grok-4.20-0309-reasoning` |
1119 | `google` | `gemini-3.1-pro-preview`, `gemini-3-pro-preview`, `gemini-3.7-flash`, `gemini-3.6-flash`, `gemini-3.5-flash`, `gemini-3.5-flash-lite`, `gemini-2.5-pro`, `gemini-2.5-flash` | yes | yes except `gemini-3.5-flash-lite` |
1120 | `mistral` | `mistral-code-latest`, `mistral-medium-latest`, `mistral-small-latest`, `mistral-large-latest` | yes | yes for Medium and Small (`reasoning_effort` `none` or `high` on exact first-party routes); deprecated native Magistral remains an always-on explicit compatibility ID; no for Code and Large |
1121 | `modelstudio-token-plan`, `modelstudio-coding-plan` | `qwen3.8-max`, `qwen3.8-max-preview`, `qwen3.7-plus`, `qwen3.7-max`, `qwen3.6-flash`, `deepseek-v4-pro`, `deepseek-v4-flash-0731`, `glm-5.2` | yes | yes |
1122 | `csdn` | `glm_for_coding`; other marketplace model IDs pass through | yes | yes |
1123
1124 AtlasCloud keeps the same default model as the config layer and adds
1125 provider-scoped aliases for the Pro and Flash rows. Other AtlasCloud model IDs
1126 should still be selected through `ATLASCLOUD_MODEL`, config, or live model
1127 listing when available.
1128
1129 ## Capability Metadata
1130
1131 `codewhale doctor --json` exposes the `capability` object. It is static
1132 metadata, not a live API probe. Current fields are:
1133
1134 `resolved_provider`, `resolved_model`, `context_window`, `max_output`,
1135 `thinking_supported`, `cache_telemetry_supported`, and `request_payload_mode`.
1136
1137 When configuration cannot be loaded or validated, `doctor --json` exits
1138 nonzero and prints a bounded, secret-redacted JSON error envelope with
1139 `status = "error"` and `error.kind = "config_validation"` instead of emitting
1140 misleading route or capability metadata.
1141
1142 Most shipped providers use the Chat Completions request payload mode. Native
1143 Messages routes, including `minimax-anthropic`, use `/v1/messages`, and
1144 `openai-codex` uses Responses.
1145
1146 For OpenAI-compatible gateways or self-hosted runtimes whose real window
1147 differs from the static table, set `[providers.<name>] context_window = N`.
1148 The configured value becomes the route-effective context window for prompts,
1149 context-pressure checks, compaction, and output-cap budgeting.
1150
1151 `max_output` is optional and truthful: it is `null` (and omitted from the
1152 capability struct on the wire) when the route publishes no output maximum we
1153 can stand behind — the Kimi Code membership `kimi-for-coding` family is the
1154 canonical example, since the membership catalog owns their limits. An unknown
1155 output ceiling is never backfilled with a placeholder, and it applies **no**
1156 compatibility clamp to a turn's requested `max_tokens`; only a concrete
1157 route/offering maximum narrows the request. A model the catalogue simply has no
1158 row for is a different fact — absence is not permission, so an uncatalogued id
1159 keeps a conservative ceiling. The "Max output metadata" column below reads
1160 `unknown` wherever no documented maximum exists.
1161
1162 The exact Kimi Code membership roster contains `k3`, `k3-256k`,
1163 `kimi-for-coding`, and `kimi-for-coding-highspeed`. The two K3 ids share the
1164 same reasoning and fixed-sampling contract; `k3-256k` stays at 262,144 tokens,
1165 while bare `k3` can use an entitled 1M override.
1166
1167 | Provider/model class | Context window | Max output metadata | Thinking support | Cache telemetry | FIM endpoint |
1168 | --- | --- | --- | --- | --- | --- |
1169 | DeepSeek V4 (`deepseek-v4-pro`, `deepseek-v4-flash`) | 1,000,000 | 384,000 | yes | yes | DeepSeek beta only |
1170 | DeepSeek V4 Flash Vision experimental (`deepseek-v4-flash-vision-exp`) | 1,000,000 inherited from Flash | 384,000 inherited from Flash | yes, inherited | yes, inherited | not claimed; Chat Completions route only |
1171 | DeepSeek compatibility aliases (`deepseek-chat`, `deepseek-reasoner`) | 1,000,000 | 384,000 | yes | yes | DeepSeek beta only |
1172 | NVIDIA NIM V4 registry models | 1,000,000 | 384,000 | yes | yes | not documented in code |
1173 | Volcengine Ark V4 model IDs | 1,000,000 | 384,000 | yes | yes | not documented in code |
1174 | OpenRouter, Novita, Fireworks, SiliconFlow, SGLang, and vLLM V4 model IDs | 1,000,000 | 384,000 | yes | no | not documented in code |
1175 | Xiaomi MiMo `mimo-v2.5-pro`, `mimo-v2.5-pro-ultraspeed`, `mimo-v2.5` | 1,000,000 | 131,072 | yes | no | not documented in code |
1176 | OpenRouter Qwen 3.6 Flash / Plus | 1,000,000 | 65,536 | yes | no | not documented in code |
1177 | OpenRouter Qwen 3.6 35B / 27B | 262,144 | 262,140 | yes | no | not documented in code |
1178 | OpenRouter Qwen 3.6 Max Preview | 262,144 | 65,536 | yes | no | not documented in code |
1179 | OpenAI API `gpt-5.5` | 1,050,000 | 128,000 | yes | no | not documented in code |
1180 | OpenAI API `gpt-5.6`, `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna` | 1,050,000 | 128,000 | yes | no | not documented in code |
1181 | Anthropic API `claude-opus-5`, `claude-opus-4-8`, `claude-sonnet-5`, `claude-sonnet-4-6`, `claude-fable-5` | 1,000,000 | 128,000 | yes | yes | not documented in code |
1182 | Google Gemini API `gemini-3.7-flash`, `gemini-3.6-flash`, `gemini-3.5-flash`, `gemini-3.5-flash-lite`, `gemini-3.1-pro-preview`, `gemini-2.5-pro`, `gemini-2.5-flash` | 1,048,576 | 65,536 | model-dependent | no | not documented in code |
1183 | Meta Model API `muse-spark-1.2` | 1,000,000 | 32,000 | yes | no | not documented in code |
1184 | ChatGPT plan route (`openai-codex`) | conservatively budgeted; account listing does not state a window | no output cap sent in preview | model dependent | no | public `/v1/responses`; no context limit is inferred from account eligibility |
1185 | OpenModel default/custom model IDs | 200,000 fallback unless model metadata or config overrides it | 64,000 fallback | model-dependent | no | route uses Messages payload at `/v1/messages` |
1186 | Wanjie Ark `reasoner` / `r1` model IDs | 128,000 | unknown (no documented maximum) | yes | no | not documented in code |
1187 | Direct Arcee API `trinity-large-thinking` | 262,144 | 262,144 | yes | no | not documented in code |
1188 | Direct Arcee API `trinity-large-preview` | 262,144 | unknown (no documented maximum) | no in doctor capability metadata | no | not documented in code |
1189 | Direct Moonshot `kimi-k3` | 1,048,576 | 1,048,576 documented maximum; 131,072 provider default | yes | no | exact route uses `max_completion_tokens` and omits fixed sampling fields ([K3 quickstart](https://platform.kimi.ai/docs/guide/kimi-k3-quickstart)) |
1190 | Kimi Code membership `k3` | 262,144 safe baseline; 1,048,576 with an explicit entitled-plan override | 131,072 conservative default ceiling; membership maximum is not published | yes | no | exact `https://api.kimi.com/coding/v1` route |
1191 | Kimi Code membership `k3-256k` | 262,144 fixed | 131,072 conservative default ceiling; membership maximum is not published | yes | no | exact `https://api.kimi.com/coding/v1` route |
1192 | Direct Moonshot/Kimi K2.7/K2.6 (`kimi-k2.7-code`, `kimi-k2.7-code-highspeed`, `kimi-k2.6`) | 262,144 | 32,768 | yes | no | provider-reported bundled catalog |
1193 | Kimi Code membership `kimi-for-coding`, `kimi-for-coding-highspeed` | 262,144 | unknown — the membership catalog owns these limits and no client-side ceiling is claimed | yes | no | exact `https://api.kimi.com/coding/v1` route |
1194 | Direct Z.AI `GLM-5.3` (default) | 1,000,000 | 131,072 | yes | no | live on the GLM Coding Plan; limits inherited from `GLM-5.2` until Z.ai publishes distinct 5.3 numbers; no USD price |
1195 | Direct Z.AI `GLM-5.3-Flash` | 1,000,000 | 131,072 | yes | no | natively multimodal; $0.15/$0.50 list (2026-08-26); faster/explore sibling of `GLM-5.3` |
1196 | Direct Z.AI `GLM-5.2` | 1,000,000 | 131,072 | yes | no | not documented in code |
1197 | Direct Z.AI `GLM-5.1` | 202,752 | 131,072 | yes | no | not documented in code |
1198 | Direct Z.AI `GLM-5-Turbo` | 202,752 | 131,072 | yes | no | faster/explore sub-agent sibling |
1199 | Direct MiniMax `MiniMax-M3` | 1,000,000 | 524,288 | yes | no | not documented in code |
1200 | Direct MiniMax M2.x models | 204,800 | unknown until MiniMax output metadata is promoted | yes | no | not documented in code |
1201 | MiniMax Messages route (`MiniMax-M3`, `MiniMax-M2.7`) | model-specific values above | model-specific values above | yes | no | route uses `/anthropic/v1/messages` |
1202 | Generic `openai` and AtlasCloud | 128,000 | unknown (no documented maximum) | no in doctor capability metadata | no | not documented in code |
1203 | Ollama | 8,192 | unknown (no documented maximum) | no | no | not documented in code |
1204 | Hugging Face Inference Providers V4 model IDs | 131,072 | unknown (no documented maximum) | yes | no | not documented in code |
1205 | Other recognized DeepSeek model IDs | 128,000 unless the model name carries an explicit `Nk` hint | unknown (no documented maximum) | no unless V4/reasoner logic matches | DeepSeek/NIM only | DeepSeek beta only |
1206
1207 MiniMax M3 uses input-length and service tiers. Codewhale omits
1208 `service_tier`, so requests use the standard tier and cost estimates select the
1209 correct standard rate from total input usage. Priority rates are listed to keep
1210 the official tier structure visible. Prices are USD per million tokens.
1211
1212 | Model / service tier | Input length | Input | Output | Cache read | Cache write |
1213 | --- | --- | ---: | ---: | ---: | ---: |
1214 | `MiniMax-M3` standard | up to 512,000 input tokens | $0.30 | $1.20 | $0.06 | not published |
1215 | `MiniMax-M3` standard | over 512,000 input tokens | $0.60 | $2.40 | $0.12 | not published |
1216 | `MiniMax-M3` priority | up to 512,000 input tokens | $0.45 | $1.80 | $0.09 | not published |
1217 | `MiniMax-M3` priority | over 512,000 input tokens | $0.90 | $3.60 | $0.18 | not published |
1218 | `MiniMax-M2.7` standard | all supported inputs | $0.30 | $1.20 | $0.06 | $0.375 |
1219
1220 These values come from the [MiniMax pay-as-you-go pricing
1221 guide](https://platform.minimax.io/docs/guides/pricing-paygo). M3 thinking is
1222 adaptive or disabled; the OpenAI-compatible API defaults to adaptive and the
1223 Anthropic-compatible API defaults to disabled. M2.7 thinking cannot be
1224 disabled. Codewhale sends explicit controls when the user selects a reasoning
1225 mode.
1226
1227 Tool-call support is tracked separately by the static `ModelRegistry` and by
1228 the endpoint's ability to accept OpenAI-compatible `tools` payloads. A custom
1229 OpenAI-compatible or local endpoint can still reject tool calls even if
1230 Codewhale can send the schema.
1231
1232 ### Hugging Face Inference Providers Notes
1233
1234 The shipped Hugging Face route targets the OpenAI-compatible Inference Providers
1235 router at `https://router.huggingface.co/v1`. Configure auth with
1236 `HUGGINGFACE_API_KEY` first, or `HF_TOKEN` as a fallback. Configure the endpoint
1237 with `HUGGINGFACE_BASE_URL` first, or `HF_BASE_URL` as a fallback; configure the
1238 model with `HUGGINGFACE_MODEL` first, or `HF_MODEL` as a fallback.
1239
1240 This route does not imply Hub browsing, model-card metadata, dataset access,
1241 Jobs, uploads, or export. Those remain explicit Model Lab work items so
1242 provider auth and artifact movement stay separate.
1243
1244 ### When a Local Model Prints Tool JSON
1245
1246 Codewhale only executes tools when the provider returns Chat Completions
1247 `tool_calls` or streamed `delta.tool_calls`. If a local model prints text such
1248 as `{"name":"File","arguments":{"action":"search_content",...}}` in the
1249 assistant message, that is ordinary model output, not an executable tool
1250 request.
1251
1252 For OpenAI-compatible or local runtimes, check:
1253
1254 - The endpoint accepts the `tools` array in `/v1/chat/completions` requests.
1255 - The selected model or chat template is configured for function/tool calls.
1256 - The server returns `tool_calls` in the response rather than plain JSON text.
1257 - The compatibility layer does not strip tools before forwarding the request.
1258 - If in doubt, test a small `File` `read` or `search_content` action against a
1259 known tool-calling model before debugging Codewhale's tool registry.
1260
1261 Changing `provider`, `base_url`, or `model` can select a route that supports the
1262 OpenAI-compatible payload shape, but Codewhale cannot convert arbitrary JSON
1263 text into a trusted tool call after the model has emitted it as prose.
1264
1265 DeepSeek will retire `deepseek-chat` and `deepseek-reasoner` on 2026-07-24 at
1266 15:59 UTC. Codewhale migrates either name to `deepseek-v4-flash` before a
1267 request reaches DeepSeek's first-party OpenAI or Anthropic endpoint. If no
1268 reasoning tier was configured, `deepseek-chat` also migrates to `off` and
1269 `deepseek-reasoner` to `high`, preserving their former non-thinking / thinking
1270 intent; an explicit `reasoning_effort` remains authoritative. The mapping is
1271 deliberately not global: Wanjie Ark, aggregators, self-hosted runtimes, and
1272 custom endpoints continue to own their model ids.
1273
1274 ## Reasoning Effort
1275
1276 `/reasoning <effort>` (and the `reasoning_effort` config key) is translated to
1277 each provider's wire dialect by the client before the request is sent. `off`
1278 disables thinking where the route supports it. Both exact K3 routes map `off`
1279 to their lowest supported tier, `low`, and the model is never switched to
1280 satisfy `off` — but they do so for different reasons:
1281
1282 - **Kimi Code membership K3** (exact `https://api.kimi.com/coding/v1` with
1283 `model = "k3"` or `model = "k3-256k"`) — the membership roster declares K3 always-thinking, so `off`
1284 cannot be honored without changing what the model is. The clamp preserves the
1285 fixed K3 identity.
1286 - **Direct Moonshot K3** (exact `https://api.moonshot.ai/v1` with
1287 `model = "kimi-k3"`) — this clamp is *defensive*, not a documented contract.
1288 The direct platform publishes no `off` state for K3, and Codewhale will not
1289 assert a fixed-thinking guarantee it cannot verify for a given key's
1290 entitlement, so the requested `off` is normalized to the lowest tier with the
1291 live entitlement left unknown.
1292
1293 Normal dispatched
1294 `auto` uses Codewhale's auto-reasoning selector and sends a concrete tier;
1295 only an omitted reasoning setting leaves the provider default in control.
1296 Providers marked "omitted" receive no reasoning fields at all for that tier.
1297
1298 | Provider | `off` | `low`/`medium`/`high` | `max`/`xhigh` |
1299 | --- | --- | --- | --- |
1300 | `deepseek`, `deepseek-cn`, `siliconflow`, `siliconflow-CN`, `sglang`, `volcengine`, `atlascloud` | `thinking: {type: disabled}` | `reasoning_effort: "high"` + `thinking: {type: enabled}` | `reasoning_effort: "max"` + `thinking: {type: enabled}` |
1301 | `openrouter`, `novita`, other `together` models | `thinking: {type: disabled}` | `reasoning_effort` pass-through + `thinking: {type: enabled}` | `reasoning_effort: "xhigh"` + `thinking: {type: enabled}` |
1302 | `together` + `thinkingmachines/inkling` | `reasoning_effort: "none"` | exact `minimal`/`low`/`medium`/`high` `reasoning_effort` | `reasoning_effort: "max"` |
1303 | Direct Moonshot `kimi-k3` at exact `https://api.moonshot.ai/v1` | top-level `reasoning_effort: "low"` (effective normalization) | top-level `reasoning_effort: "low"` / `"high"` (`medium` becomes `high`) | top-level `reasoning_effort: "max"` |
1304 | Kimi Code membership `k3`, `k3-256k` at exact `https://api.kimi.com/coding/v1` | `thinking: {type: enabled, effort: "low"}` (effective normalization) | `thinking: {type: enabled, effort: "low" | "high"}` | `thinking: {type: enabled, effort: "max"}` |
1305 | Other `moonshot` routes | `thinking: {type: disabled}` | `thinking: {type: enabled}` | `thinking: {type: enabled}` |
1306 | `ollama` | `think: false` | `think: true` | `think: true` |
1307 | `ollama-cloud` | `reasoning_effort: "none"` | exact `low`/`medium`/`high` `reasoning_effort` | `reasoning_effort: "max"` |
1308 | `xiaomi-mimo` | `thinking: {type: disabled}` | `thinking: {type: enabled}` | `thinking: {type: enabled}` |
1309 | First-party `minimax` `MiniMax-M3` | `reasoning_split: true` + `thinking: {type: disabled}` | `reasoning_split: true` + `thinking: {type: adaptive}`; effective tier granularity unavailable | `reasoning_split: true` + `thinking: {type: adaptive}`; effective tier granularity unavailable |
1310 | First-party Z.ai `GLM-5.2` | `thinking: {type: disabled}`; no `reasoning_effort` | enabled thinking; only effective `high` adds `reasoning_effort: "high"` | enabled thinking + `reasoning_effort: "max"` |
1311 | First-party Z.ai `GLM-5.3` | `thinking: {type: disabled}`; no `reasoning_effort` | enabled thinking; only effective `high` adds `reasoning_effort: "high"` | enabled thinking + `reasoning_effort: "max"` |
1312 | First-party Z.ai `GLM-5.3-Flash` | `thinking: {type: disabled}`; no `reasoning_effort` | enabled thinking; only effective `high` adds `reasoning_effort: "high"` | enabled thinking + `reasoning_effort: "max"` |
1313 | First-party Z.ai `GLM-5-Turbo` | `thinking: {type: disabled}` | enabled thinking; effort granularity unavailable | enabled thinking; effort granularity unavailable |
1314 | Compatible gateways configured as `zai` | omitted; effective unavailable | omitted; effective unavailable | omitted; effective unavailable |
1315 | `nvidia-nim` | `chat_template_kwargs.thinking: false` | `chat_template_kwargs`: `thinking: true` + `reasoning_effort: "high"` | `chat_template_kwargs`: `thinking: true` + `reasoning_effort: "max"` |
1316 | `vllm` | `chat_template_kwargs.enable_thinking: false` | `chat_template_kwargs.enable_thinking: true` + `reasoning_effort` low/medium/high | `chat_template_kwargs.enable_thinking: true` + `reasoning_effort: "high"` (vLLM has no max tier) |
1317 | `arcee`, `huggingface` | omitted | `reasoning_effort` pass-through | `reasoning_effort: "high"` |
1318 | `fireworks` | omitted | `reasoning_effort: "high"` | `reasoning_effort: "max"` |
1319 | `openai`, `wanjie-ark`, `telecomjs` | omitted | omitted | omitted |
1320 | `openmodel` | Anthropic Messages adapter handles thinking/output configuration | Anthropic Messages adapter handles thinking/output configuration | Anthropic Messages adapter handles thinking/output configuration |
1321 | `openai-codex` | Responses API `reasoning` field (handled by the Responses bridge) | Responses API `reasoning` field | Responses API `reasoning` field |
1322
1323 AtlasCloud serves DeepSeek models, so it speaks the DeepSeek reasoning dialect,
1324 including the `max` tier (#3024).
1325
1326 On the exact MiniMax OpenAI-compatible Chat endpoints, `MiniMax-M3` uses
1327 `max_completion_tokens`. Other MiniMax models and compatible gateways retain
1328 `max_tokens`; the MiniMax Anthropic endpoints use the separate Messages
1329 adapter.
1330
1331 ## Drift Check
1332
1333 Run this before changing provider IDs, provider TOML tables, static model
1334 registry rows, or provider default strings:
1335
1336 ```bash
1337 python3 scripts/check-provider-registry.py
1338 ```
1339
1340 The check fails when:
1341
1342 - `docs/PROVIDERS.md` omits a canonical `ProviderKind::as_str()` ID.
1343 - Descriptor presentation IDs or released wire tags are duplicated or drift
1344 from intrinsic kinds, or a second provider enum/ordinal bridge is introduced.
1345 - The shipped-provider table omits or adds a `[providers.*]` TOML table.
1346 - The static model registry table drifts from providers used by
1347 `crates/agent/src/lib.rs`.
1348 - A provider default model or base URL constant in `crates/tui/src/config.rs`
1349 is no longer mentioned here.
1350
1351 ## Planned, Not Shipped Yet
1352
1353 These items belong to the v0.8.48+ provider-abstraction milestone or related
1354 provider docs work, but they are not native shipped behavior in this checkout:
1355
1356 - A unified `Provider` trait in `codewhale-agent` that owns env precedence,
1357 secret resolution, base URL normalization, auth-header construction, and
1358 provider metadata. Those responsibilities are still split across
1359 `crates/config`, `crates/secrets`, and `crates/tui/src/client.rs`.
1360 - Hugging Face model passport metadata in the picker, including license, base
1361 model, context length, chat template, tool-call support, reasoning support,
1362 and gated/private status.
1363
1363 lines MARKDOWN