Get started
v2026.8.1: Models and Providers
Chat and the Models page now open from the catalog OpenClaw already has, so choosing a model no longer waits for a full provider scan. Live discovery runs only when you open a model screen or explicitly refresh it, keeps the last useful list when a lookup fails, and /model can change only the current conversation or deliberately update one agent or the shared default.
Once a model is selected, OpenClaw keeps the request inside the intended provider and authorized account order, preserves the real order and outcome of streamed replies and tool calls, carries reasoning and context settings with the model and runtime, and reports plan windows, token use, context pressure, and estimated cost more clearly.
Finding and choosing models
Chat and the Models page now start from the catalog OpenClaw already has, with supported providers looking for newer chat and text models when you open a picker or request a refresh. If that lookup fails, the built-in entries and last working list remain available.
A /model change can stay with the current conversation or deliberately apply to one agent or the shared default, with persistent changes requiring the right authority. Aliases and fallback keep the provider and account attached to the selected model.
The list shows models OpenClaw can identify, while the provider, account, region, endpoint, plan, limits, and pricing determine which ones you can use.
Sources and complete change list
Improvements
- Migrate legacy Codex model routes into OpenAI 3c8269c
- Show provider-configured models on Android #101912
- Add Grok 4.5 to the bundled xAI provider #102316
- Refresh Qwen, Cohere, and Mistral model catalogs #102489
- Add current Cohere Command models #102563
- Add bundled Meta Model API support for Muse Spark 1.1 #102873
- Publish Meta under its canonical provider package #103070
- Derive utility models from the primary provider #103769
- Move active defaults and examples to GPT-5.6 Sol and Luna #104452
- Reduce repeated agent turn setup latency #105220
- Add official Baseten Model API support #108708
- Add Kimi K3 support across Moonshot and Kimi Code #109202
- Add provider-grouped models and per-session controls to macOS chat #110339
- Discover current chat models from live provider catalogs #112412
- Make provider catalogs authoritative for model compatibility #112542
- Add end-to-end Claude Opus 5 support #113391
- Complete Claude Opus 5 support across Anthropic routes #113392
- Curate model pickers around current-generation choices #113594
- Complete Claude Opus 5 support across Anthropic routes #113633
- Refresh model catalogs without an OpenClaw release #113660
- Refresh provider catalogs to current model lineups #113681
- Start warm agent turns faster #113817
- Add first-class Kimi K3 support #113909
- Consolidate model settings on one Models page #115108
- Follow xAI's OAuth default model automatically #115617
- Clarify Model Setup and remove duplicate active routes #116086
- Add session-only model selection and preserve fallback state #119325
- Add Meta Muse Spark 1.2 models #120373
- Add first-class Grok 4.6 catalog and reasoning support #122762
- Add GLM 5.3 support for Z.AI Coding Plan #123523
- Add first-class support for existing llama-server endpoints #125781
- Unify managed and existing llama.cpp servers #126434
- feat(models): add configurable model selection scopes #127813
- feat(models): refresh provider models and live discovery #131294
- Add first-class Alibaba Qwen Token Plan support #94419
- Update the bundled Codex app-server to 0.144.1 #103156
- Reduce warm Gateway model startup latency #105552
- Refresh Z.AI, Kimi, Moonshot, and xAI catalog guidance #109666
- Show bundled provider icons during model setup #112169
- Speed up provider-filtered model listing #112752
- Refresh OpenCode Zen models and pricing #113858
- Add the Kimi K3 256K model option #113918
- Keep code-mode tiers consistent across provider catalogs #115183
- Remember the last-used model for future sessions #115717
- Show full model names alongside aliases #115789
- Clarify local model setup actions #116953
- Reduce memory use for the default model list #117323
- Clarify llama.cpp local model setup #117947
- Speed up verified model setup #118352
- Remove full model-catalog builds from turn startup #120834
- Explain the Utility model setting in Models #123277
- Reuse the WebUI model catalog across pages #124794
- Avoid repeated plugin scans during model lookups #125090
- perf(tui): prepare searchable picker rows once #127620
- fix(ui): add brand icons for all provider auth choices #127668
- perf(cli): lazy-load infer command domains #128150
- feat(zai): add GLM-5.3 Flash model support #130309
- Add Step 3.7 Flash to StepFun model catalogs #88082
- Reuse prepared plugin metadata in static model catalogs #117288
- Show model tables without waiting on promotions #117322
- Avoid full Ollama plugin loading during model listing #117465
- perf(models): reuse projected list rows #128248
Bug fixes
- Refresh xAI models, reasoning, and server-tool defaults #103260
- Align OpenAI model availability with its effective route #104685
- Harden OpenAI Codex discovery, opt-out, and quota recovery #108683
- Refresh bundled provider catalogs and remove retired models #109410
- Separate model aliases from override allowlists #110888
- Apply current request shaping to future Claude generations #114205
- Unify dynamic model catalog selection #114284
- Preserve selected fallback providers across alias resolution #114360
- Keep model lists consistent across CLI and Gateway #114393
- Keep agent model catalogs and switching reliable #114760
- Improve gateway-node and local-model reliability #115185
- Prevent multi-agent gateway startup stalls #116261
- Stop cloud fallback after stale Gateway lifecycle aborts #117168
- Keep large-roster chat startup off model catalog discovery #119742
- Fall back when a model returns no usable answer #120148
- Keep new-chat inference available across routes #120712
- Unify agent failover error classification #121341
- Keep model catalog discovery from stalling the Gateway #122350
- Keep session model selection separate from response aliases and fallbacks #122910
- Restore model lists when chat still works #123208
- Keep Codex-routed turns working with request overrides #123258
- Prevent multi-agent Gateway startup CPU spikes #124334
- Let direct inference commands select an agent #125143
- Honor per-agent model metadata everywhere #126194
- fix(opencode): discover Zen and Go models on demand #129831
- fix(models): keep native and alias-backed catalogs consistent #131677
- fix: allow off-catalog chat models under unrestricted policy #132175
- Keep collected follow-ups on their model routes 1647027
- Keep model fallbacks working after cleanup takeovers #100118
- Add Claude Fable 5 to the Claude CLI model catalog #101453
- Expose normal model fallback transitions in diagnostics #102051
- Apply hot-reloaded agent models without restarting the gateway #102305
- Match Azure deployment map IDs across casing differences #103005
- Keep replace-mode model listings limited to configured providers #103103
- Restore current OpenCode Zen models in offline discovery #103197
- Support managed ClawRouter gateway configuration #103299
- Remove deprecated OpenCode Go MiMo aliases #103329
- Preserve Grok behavior through the x-ai alias #103340
- Allow the Codex runtime with
codex/*models #103775 - Surface utility-model narration failures in logs and model status #104770
- Bound Amazon Bedrock model discovery requests #104841
- Explain account-restricted model failures correctly #104878
- Keep locked sessions on their stored model route #105335
- Keep native thinking and fast controls on Codex routes #107588
- Preserve newer model and auth choices during fallback recovery #108137
- Bound Amazon Bedrock control-plane requests #108756
- Bound stalled Z.AI endpoint probe responses #109026
- Release Hugging Face connections after failed model discovery #109464
- Preserve pinned models during temporary catalog failures #110364
- Add OpenCode Zen GPT-5.6 model discovery #110405
- Report unavailable Codex runtimes in model status #110447
- Recover network error codes for model failover #110485
- Report accurate OpenRouter limits in model scans #110855
- Discover Vercel AI Gateway models through proxies #111209
- Hide image-output models from the Kilocode chat catalog #111239
- Discover Chutes models through eligible HTTP proxies #111266
- Reuse model discovery across configless embedded runs #111423
- Reject unknown text models before saving config #111571
- Give cold model-runtime builds a longer startup budget #111983
- Keep model catalogs coherent during config reloads #112331
- Honor explicit subagent model requests #112393
- Detect Claude CLI routes pinned to older models #113424
- Hide live-discovered Claude models until request contracts support them #113757
- Fix Anthropic live discovery for Claude subscription users #113906
- Keep concurrent model status JSON clean #113953
- Reduce reply startup delay with configured model policies #114117
- fix(agents): retry live model switches outside fallback chain #114214
- Show refreshed OpenAI models in provider-filtered lists #114265
- Fall back on raw provider request errors #114457
- Stabilize model metadata and config recovery under load #114679
- Repair malformed generated model catalogs before startup #114697
- Recognize local model endpoints and cap fallback output #114710
- Keep Infer provider inspection and selection working #115054
- Stabilize infer model selection, provider status, and shared-state access #115170
- Show correct plugin-catalog metadata in configured model lists #115190
- Canonicalize model references at their input owners #115965
- Prevent large model configurations from delaying Gateway startup #116553
- Show local provider identity after model setup #116655
- Keep healthy fallback models available after a missing-model error #116768
- Restore runtime-discovered models in browse views #116857
- Reuse plugin metadata when building large model catalogs #117131
- Keep alternate runtimes available after Codex preflight failures #117256
- Use configured fallbacks after replay-safe silent failures #117356
- fix(tui): honor restrictive agent model policies #117422
- Restore Cerebras's documented static model catalog #117584
- Hide retired NVIDIA models while preserving live discovery #117620
- Prevent model discovery from hanging past its timeout #117825
- Preserve model selection while the gateway is draining #118101
- Preserve newer plugin metadata during models status #119183
- Honor configured models for ACP agents #120046
- Refresh OpenCode Zen and Go model catalogs #121030
- Make exact dated model pins executable after Gateway startup #121119
- Migrate retired OpenCode free-model references #121125
- Show accurate defaults while a new session loads #121133
- Deliver replies after CLI-to-embedded model fallback #121897
- Show the resolved default model and its thinking levels #122907
- Show resolved model identities in agents_list #123044
- Keep scoped model checks from building the full catalog #123473
- Stop Models page agent-roster retry storms #123483
- Prevent truncated verified llama.cpp downloads #123738
- Scope inference provider lists to a selected agent #123884
- Let older llama.cpp providers load for upgrade repair #124041
- Keep session status accurate through model fallback #124436
- Preserve plugin completion agent selection on multi-agent installs #124954
- Reject unknown providers when selecting default models #125558
- Keep in-flight model requests on one plugin generation #125569
- fix(providers): keep discovered image support for configured models #125670
- Keep inherited parent model pins ahead of channel defaults #125880
- Reject unusable explicit subagent model requests #125931
- Preserve catalog capabilities for partial provider model configuration #126068
- Scope local inference preparation to the selected model #126196
- Let multi-agent infer commands select their owning agent #126234
- Preserve explicit Codex runtime policy for embedded attempts #126259
- fix(llm-task): trust explicit provider for slash-containing model ids #126518
- Preserve Codex selection when model metadata comes from the catalog #126531
- refactor(providers): return prepared dynamic models directly #126574
- fix(codex): run newly available models through the existing Codex account #127322
- Let Codex use newly available account models before catalog refresh #127394
- fix(models): scope status alias resolution to the selected agent #127631
- fix(models): stop /models from freezing the gateway while it rescans plugins #127847
- Skip inherited automatic fallback overrides in child sessions #128521
- perf(gateway): stop first models.authStatus call from reloading all plugins from source #128551
- fix(workers): resolve catalog models on dispatched turns #128553
- fix: /models stops working after an unrelated config edit and never recovers #128608
- fix(tui): restore session setting reset directives #128731
- fix(codex): keep /codex model scoped to the native binding #128751
- fix(vercel): exclude non-language gateway models from chat catalog #129406
- fix: agent-specific model selection repeatedly rescans every plugin #129510
- fix(huggingface): honor available provider capabilities #129797
- fix(cli): preserve human output for flag-like model option values #129911
- fix(models): model browsing fails after automatic plugin activation #130481
- fix(agents): preserve Codex for reasoning capability metadata #130752
- fix(models): keep OpenCode fallback aligned with catalog lifecycle #130785
- fix(models): keep model listings consistent after manifest changes #131518
- fix(config): accept supported model pins outside curated catalogs #131562
- fix(codex): keep native overrides effective at launch and turn start #131745
- fix(agents): preserve native runtime auth across catalog refresh #131931
- fix(agents): retain prepared catalog ownership through refresh #132244
- fix(anthropic): defer Claude CLI capability probe until execution #132324
- fix(fireworks): prevent unsupported images reaching GLM 5.2 Fast #132362
- perf(plugins): consolidate provider discovery planning #133097
- Recognize code-only provider errors for model failover #95729
- Fix Codex ACP spawns with non-OpenAI fleet defaults #95852
- Keep generated multimodal models in the registry #97858
- Refresh NVIDIA featured models without breaking older configs #99242
- Align Codex model discovery with client version 0.144.5 cfaa30a
- Keep truncated fallback reasons valid Unicode #102631
- Keep model tables aligned with wide Unicode names #102819
- Align Meta validation and guidance with Muse Spark 1.1 #103163
- Preserve Unicode in bounded OpenRouter Fusion model IDs #104433
- Preserve Codex mock routing with model-specific image endpoints #105444
- Repair legacy Codex model references in per-agent model maps #108369
- Allow clearing agent model overrides #108734
- Honor OpenRouter scan timeout while reading the catalog #108972
- Release stalled GitHub Copilot catalog error streams #109071
- fix(plugin-sdk): contextualize malformed provider catalogs #109754
- Handle malformed provider catalog pagination URLs safely #109986
- Make status reflect pending and completed live model switches #110474
- Discover Hugging Face models through required HTTP proxies #110924
- Sync OpenAI Codex OAuth catalog client version #110936
- Reject malformed UTF-8 during self-hosted model discovery #111113
- Nested model-policy wildcards respect namespace boundaries #111350
- Let DeepInfra model discovery use trusted HTTP proxies #111428
- Reject corrupted UTF-8 in live provider catalogs #111766
- Keep new chats working through concurrent config reloads #112026
- Report model switches before automatic thinking remaps #112352
- Preserve local-provider errors during cron preflight #113409
- Align Codex OAuth model discovery with the managed client version #113615
- Prevent Meta catalog loading from rejecting unsupported inputs #113758
- Keep empty model lists machine-readable #113897
- Accept harmless separator padding in model policy references #114116
- Make refreshed model catalogs usable from the CLI #114244
- Recognize all IPv4 loopback model endpoints #114766
- Reject incomplete model overrides in infer commands #115225
- Align configured model rows with catalog availability #115239
- Preserve valid LM Studio models in mixed catalogs #115336
- fix(status): recent session listing resolves the wrong agent's model #115984
- Reject malformed hosted model catalogs before saving #117744
- Preserve explicit provider and model references in alias guidance #117869
- Report accurate embedded-agent startup timing stages #117879
- fix(config): ModelCompatSchema drops eight keys ModelCompatConfig declares, and the drift contract cannot see it #118674
- Apply the global fast-mode default to the implicit agent #120697
- fix(acpx): strip amazon-bedrock provider prefix from Claude ACP model refs #121105
- Preserve concrete provider failure reasons through failover #121657
- Propagate provider catalog failures instead of mistaking them for timeouts #124288
- Preserve complete model fallback traces #126048
- Stop stale provider tests after switching agents #126750
- fix(models): --agent is silently ignored by models aliases and models scan #126864
- fix(cli): return JSON for model list failures #127726
- fix(agents): resolve pdf tool models through the canonical resolver #127952
- fix(agents): warn when configured image model entries fail to resolve #127971
- fix(cli): resolve alias targets from the write snapshot #128223
- fix(cli): render models refresh JSON failures #128981
- fix(cli): keep models --plain stdout clean of startup diagnostics #129037
- fix(models): normalize mixed aliases consistently #129678
- fix(agents): prevent selected-agent alias stalls #129700
- fix(reply): keep fast abort off provider runtime #129704
- fix(tui): dismiss 'loading models...' notice once the model picker opens #129770
- fix(models): avoid retired automatic auth-probe models #129798
- fix(cli): preserve populated model output on stdout in plain mode #129871
- fix(ui): reset /model default through server directive #129895
- refactor(agents): resolve runtime models through prepared owner #130007
- fix(bedrock): reject invalid UTF-8 model discovery #130058
- fix(agents): read the published model catalog owner from session_status #130277
- refactor(models): preserve catalog metadata through one normalization path #130381
- fix(cli): reject --agent for global model refresh #130448
- fix(cli): reject empty agent scopes for global model commands #130739
- fix: model picker refresh discovers configured agent catalog #131151
- fix: cold inventories incorrectly quarantine provider-supported tools #131601
- fix(config): preserve authored model rows during sparse patches #132080
- fix(codex): explicitly selected hidden models fail bounded turns #132189
- fix(sessions): honor per-agent model policies and aliases #132205
Documentation
- Clarify OpenAI and Codex Fast mode precedence #120682
- docs: correct numeric model selection guidance #132246
- Document rolling Anthropic Opus aliases #113413
- Repair OpenRouter copy-paste configuration examples #118747
- Clarify model switching and cache trade-offs #124889
- docs: correct chat model selection examples #132218
Provider accounts and sign-in
OpenClaw now keeps model requests inside the provider accounts and credential order you configured. If one account hits an authentication or quota cooldown, the next authorized account for that provider can take over without changing the selected provider or model, and the saved preference resumes when it recovers. Environment keys remain available when no explicit account list is configured.
OAuth registration stays with the setup conversation where it began, failed credential writes surface as failures, and a successful login applies only to the model route that authenticated. Tenant and custom-endpoint credentials stay on the account and origin they were configured for.
New GitHub Copilot device logins place the token in OpenClaw's protected local secret store by default and keep a reference in the auth profile. The store depends on state-directory permissions rather than encryption at rest, existing inline profiles are not migrated, and operators can still choose the prior plaintext mode explicitly.
Sources and complete change list
Improvements
- Support loopback Codex proxies in the OpenAI provider #114555
- Add GitHub Enterprise data-residency Copilot support #99221
- Remove individual model auth profiles from the CLI #114407
- Reuse SQLite readers for faster auth-profile loading #121358
- Clarify auth profile order output #112203
- Refresh macOS local-model and provider sign-in translations #113493
- Retire legacy provider IDs and secret markers at runtime #115655
Bug fixes
- Avoid false 12-hour OpenAI cooldowns after token expiry #102479
- Restore native xAI routing for the x-ai provider alias #103488
- Remove raw synthetic credentials from model status JSON #104734
- Block unsupported GitHub Copilot OAuth domains #105584
- Restore GitHub Copilot turns and account-scoped context limits #106198
- Restore service-account authentication for Anthropic Vertex #108350
- Keep OAuth providers isolated to their owning session #108680
- Surface credential saves and harden agent sessions and tools #109622
- Accept fine-grained GitHub tokens for Copilot #114282
- Honor provider auth pins for ambient credentials #114944
- Skip provider-auth scans on unrelated reloads #117449
- Preserve compatible auth profiles when switching models #117550
- Prevent undeclared environment credentials from taking over failover #118458
- Keep OpenRouter custom-proxy credentials on the configured origin #118773
- Stop retrying expired CLI auth profiles after fallback #119843
- Refresh chat model availability after runtime auth succeeds #121193
- Scope quota failures to individual auth profiles #121278
- Make Doctor converge on missing OAuth sidecars #123093
- Honor configured GitHub Copilot identity across model requests #127965
- fix(github-copilot): device login no longer leaves plaintext tokens #133033
- Refresh model availability after OAuth renewal #102289
- Redact GitHub Copilot OAuth failure details #102953
- Preserve OAuth diagnostics when provider messages contain braces #103106
- Show actionable re-auth guidance for logged-out Claude CLI #103829
- Keep Codex account identity when channel login switches profiles #104706
- Keep unresolved API-key providers visible #106754
- Stop false Claude CLI expiry warnings #108205
- Redact credentials from GitHub Copilot embedding errors #109177
- Ignore blank Google Vertex project and location values #109515
- Improve provider login, Bedrock errors, and TUI rendering #109740
- Keep Telegram responsive during Codex device login #110580
- Ignore blank AWS region overrides for Bedrock #110676
- Use canonical auth-profile ordering for background credential discovery #110678
- Release stalled GitHub Copilot OAuth error responses promptly #112268
- Use selected Claude profiles for CLI-backed agent turns #112458
- Prevent Gateway crashes after Claude CLI auth-profile turns #112550
- Retry transient OpenAI device-code polling failures #114094
- Refuse mismatched Codex native-home billing routes #114719
- Continue model fallback after Google invalid-key errors #114887
- Keep Claude CLI logins refreshable instead of forwarding stale tokens #115134
- Stop model fallback on rejected SSH sandbox secrets #115198
- Isolate coding-agent Codex authentication #115611
- Restore Copilot stored-profile authentication in external harnesses #116555
- Guide recovery when agent-scoped Codex credentials were not imported #116807
- Keep model-provider changes scoped and visible #116829
- Recover from unavailable automatic auth-profile overrides #117846
- Honor Google Cloud SDK credentials and Vertex billing projects #118745
- Recognize configured credentials for media providers #118761
- Keep transient timeouts from disabling inline API keys #119147
- Preserve inherited auth profile order for secondary agents #119505
- fix(ollama): redact reflected credentials from transport errors #119537
- Apply auth-order changes without restarting Gateway #119946
- Keep models available after Gateway restart #120977
- Preserve rate-limit retries when no fallback model is configured #121294
- Count only successful rate-limit profile rotations #121325
- fix: inherited auth-profile cooldown writes target the child store and silently no-op #122604
- Recover OpenAI sessions with a backup auth profile #123791
- Keep auth-profile rotation on the current provider route #123894
- fix(auth): reset only the profile replaced by a successful login #124014
- Deprioritize OAuth profiles after refresh failures #125507
- Restore explicit Bash environment imports without changing exec PATH probes #125624
- Report the auth database that actually holds an agent's profiles #126918
- Show the persisted auth database path #126955
- fix(agents): surface provider auth guidance for terminal failures #128375
- fix(auth): recover runtime profiles after publication failure #129234
- fix(auth): inherit the shared auth store once it owns credentials #130264
- fix(tui): reload auth store ownership after login #130349
- fix(agents): publish SecretRef Codex API-key route success #130359
- fix(models): verify native CLI auth in detailed lists #130749
- fix(github-copilot): restore usage for domain-aware OAuth profiles #131137
- Report Claude OAuth expiry with recovery steps in channel chats #131345
- fix(gateway/ui): record model auth unavailability reasons; gate composers only on actionable auth failures #131697
- fix(auth): keep working OpenAI models available after catalog refresh #131803
- fix(ui): chat stays blocked after model credentials recover #132179
- fix(process): make credential descriptors reopenable #132573
- fix(agents): separate auth profiles from runtime model ids #133175
- Cool down inline provider API keys after billing failures #88709
- Keep Azure Responses models on Azure credentials #93833
- Repair auth status and stale credential orders #98105
- Materialize auth-profile secret refs for standalone local agents d4ae2bb
- Bound Chutes OAuth requests with a 30-second deadline #102026
- Bound Google Vertex credential refreshes to 30 seconds #102050
- Keep Microsoft Foundry Azure CLI errors readable #102604
- Keep Microsoft Foundry connection errors Unicode-safe #102605
- Ignore malformed Microsoft Foundry endpoints #104796
- Show re-auth hints for inline OAuth errors #105085
- Reject malformed fallback-skip TTL values #107234
- Bound stalled OpenAI device-code login requests #109494
- Preserve Unicode in Microsoft Foundry Azure CLI login output #109499
- fix(microsoft-foundry): bound az subprocess lifetime on exec timeout #110434
- Keep OAuth refresh failures from using substitute credentials #112957
- Continue model fallback after Google invalid-key errors #114825
- fix(agents): separate retry count from API key position (#102938) #115507
- Make Telegram Codex login codes tap-to-copy #116608
- Use official Anthropic Vertex multi-region endpoints #116757
- Stop no-op OAuth migration prompts and show failure causes #123164
- Preserve Codex QA login identity across restarts #126777
- Point missing API-key errors to the auth command #127011
- fix(agents): stored auth profile order is ignored when resolving CLI runtime aliases #129165
- fix(ollama): protect local credentials on all loopback addresses #129168
- fix(auth): reject retired profiles from cold external providers #129580
- fix(ui): clear stale provider API-key drafts when switching agents #129676
- fix(github-copilot): configured account order is ignored #129978
- fix(ai): require explicit endpoints for compatible providers #130453
- fix(auth): preserve credential-source fact through model attempts #131220
- fix(auth): keep SecretRef models available with explicit auth order #131255
- Deduplicate Bedrock Mantle IAM failure diagnostics per region #81208
Documentation
Streaming Replies and Tool Calls
OpenClaw now distinguishes a complete streamed reply from one that stopped partway through. On supported Responses paths, reasoning, text, tool output, and available usage remain in order; if a stream fails, valid work that already arrived is preserved, and OpenClaw retries or moves to an authorized fallback only when replay is safe. With no safe recovery path, the partial reply remains attached to the error.
Tool calls wait for a complete name and arguments before they can run, including large streamed arguments and calls whose provider item IDs change along the way. Unmanaged native OpenAI Responses conversations on the official endpoint with storage enabled can continue by sending only new input after the first turn, then recover with full history if that upstream continuation state expires.
Sources and complete change list
Improvements
- Add structured Codex questions and native goal controls #109724
- Relay Claude CLI native tool requests through Gateway approvals #112918
- Faster OpenAI agent turns through reusable WebSockets #121687
- Continue stateful native OpenAI SSE turns #122194
- feat(codex): upgrade main to app-server 0.149.1 #128370
- Enable capability-gated tools on CLI-backed models #102835
- Prewarm Codex app server during turn setup #114551
- Reduce Codex coding-runtime and parity overhead #114574
- Stream native Ollama tool-call lifecycle events #116809
- perf(copilot): avoid duplicating BYOK request bodies #131145
- Update the managed Codex app-server to 0.146.1 b269e65
- Update the bundled Codex CLI to 0.144.4 #107208
Bug fixes
- Accept numeric and boolean enums in Gemini tool schemas #104567
- Harden Codex aborts, dynamic-tool schemas, and result middleware #108105
- Improve Codex full-auto execution, failure handling, and quota status #108626
- Preserve tool arguments when Responses item IDs rotate #108630
- Preserve tool-result text when media has no inline payload #108697
- Restore ClawRouter Perplexity tools and model refreshes #108758
- Keep tool-call preambles out of Chat Completions replies #109057
- Move ACPX Codex sessions to the maintained adapter #109075
- Harden OpenAI-compatible provider reliability and errors #109556
- Harden Responses streaming across OpenAI, Azure, and ChatGPT #109615
- Prevent repeated tool-call IDs from poisoning sessions #110518
- Enforce bounded CLI input and provider response deadlines #110627
- Unify OpenAI Responses stream processing #114263
- Restore Codex patch execution and fast-mode priority #114907
- Stop paired local inference after caller cancellation #115624
- Preserve complete streamed Chat Completions responses #115652
- Keep stressed Codex turns isolated and visible #115893
- Make Codex and OpenAI turns finish truthfully #116012
- Preserve complete terminal output in Responses streams #116836
- Repair llama.cpp tool, reasoning, and context lifecycles #116903
- Recover authoritative OpenAI Responses output #116910
- Restore reliable Gemini and Vertex streaming and usage accounting #117042
- Preserve fenced tool-call examples in replies #117089
- Preserve structured OpenAI Chat responses #117136
- Preserve Gemini signatures and provider failures #117193
- Restore OAuth-backed tools for CLI agents #119166
- Preserve the final streamed text when a run fails #119556
- Flush trailing chat text after quiet streaming gaps #119566
- Restore Venice Gemini tool-call continuations #119783
- Preserve results from quiet long-running Codex tools #119835
- Surface Claude CLI no-response failures #121589
- Deliver partial replies when models reach their output limit #123546
- Reuse Codex app servers after managed desktop fallback #124885
- Reap Codex app-server descendants during retirement #126285
- Surface Codex input prompts across native and ACP runs #126387
- Reject malformed streamed tool calls #126391
- fix(agents): retry incomplete model streams with partial output #127338
- fix(gateway): fail streaming responses when agent runs fail #127662
- fix(gateway): restore HTTP subagent completion announces [AI-assisted] #128068
- perf(ai): keep streaming responsive while large tool call arguments assemble #128166
- fix(agents): apply tool policy to Anthropic native calls #128805
- fix(agents): quiet model turns are aborted by a false "no response from model" idle timeout (#117884) 0707fcb
- fix(openai): reconcile rotated terminal message IDs (#122560) 5acfc2d
- fix(openai-responses): reconcile terminal tool calls (#108461) 8c1a55d
- Retry transient network failures consistently #101496
- Show OpenAI refusal text instead of blank replies #102344
- Restore image tool results for prefixed Gemini 2 models #102382
- Update the Codex harness for app-server 0.143.0 #102444
- Recognize zero-argument XML tool calls #103220
- Preserve streamed answer order after tool-call repair #103585
- Align the Meta provider with the official Model API #103680
- Stop canceled local-provider startups from leaking sidecars #104070
- Bound and unify serialized tool-call repair #104100
- Preserve Unicode across streamed block-chunk boundaries #104441
- Fix GPT-5.6 tool calls and catalog compatibility #104555
- Preserve explicit abort outcomes for Codex turns with unfinished tools #104955
- Clean up abort listeners after ChatGPT retry waits #105519
- Preserve newer ChatGPT Responses WebSocket connections #105634
- fix(deepinfra): apply request policy to video generation requests #105887
- Update the managed Codex app-server to 0.144.3 #106098
- Retry transient provider-wrapped 5xx errors #106850
- Report empty proxy responses clearly #107325
- Recover large Codex app-server messages #107369
- Preserve Codex assistant message ordering during image saves #107680
- Terminate Anthropic streams on circular provider errors #107800
- Correct Codex tool-result failure classification #107807
- Split embedded runner orchestration and harden timeout recovery #107900
- Return visible text from constrained Z.AI GLM completions #108612
- Recover Codex tool-terminal turns and align auth routing #108966
- fix(agents): idle-timeout stalled provider response body reads #109067
- Prevent broken Claude CLI output pipes from crashing node hosts #109794
- Stop cancelled replies from retrying after transient provider errors #109805
- Restore local GGUF asset resolution #110233
- Reject invalid HTTP status prefixes in assistant errors #110554
- Keep repeated Responses tool calls distinct #110956
- Release guarded response stream reader locks #111259
- ChatGPT Responses honor the server's retry delay fallback #111353
- Fail fast on invalid provider TLS certificates #111818
- Bound Ollama requests during stalled DNS preflight #111835
- Suppress partial assistant output after agent errors or aborts #113012
- Keep streamed OpenAI-compatible requests running after handler return #113514
- Return 404 for disabled OpenAI-compatible API routes #113609
- Close failed Responses streams without false completion #113629
- Preserve OpenAI Responses cancellations #113801
- Repair provider-polluted tool names safely #113814
- Preserve Qwen Token Plan behavior for direct model references #113976
- Unify OpenAI-compatible request behavior #114236
- Honor per-turn OpenAI timeouts and retry limits #115017
- Enable automatic code mode for Kimi Code K3 #115022
- Preserve provider timeouts across plugin model rebuilds #115102
- Preserve partial output from incomplete OpenAI Responses streams #115132
- Show Codex startup diagnostics on initialize timeout #115161
- Report embedded agent deadlines as timeouts #115319
- Enforce one timeout for paired-node Ollama inference #115558
- Accept the OpenAI SDK plain-text Responses format #115594
- Restore tool calling for llama.cpp-compatible providers #115598
- Keep progressing remote Ollama streams alive #115648
- Strip unsupported keywords from every nested tool schema #115741
- Keep OpenAI Responses replacement streams consistent #115799
- fix(gateway): preserve OpenAI-compatible response contracts #116354
- Let one-shot Claude CLI commands exit cleanly #116577
- fix(runtime): prevent workspace, provider, and streaming lifecycle leaks #116650
- Execute llama.cpp plaintext tool calls #116736
- Honor Ollama stop sequences and asynchronous payload hooks #116737
- Reject truncated Bedrock streams and preserve audio results #116743
- Surface blocked Gemini prompts with usage details #116756
- Preserve Google stream errors, usage, and tool signatures #116822
- Preserve OpenAI Chat tool-call lifecycles #116922
- fix(google): preserve stream terminals and provider tool-call identities #116926
- Show clean retry guidance for malformed Codex streams #116966
- Preserve fenced tool-marker examples in replies #116971
- Repair Bedrock stream completion and OpenRouter model handling #116998
- Enforce streamed OpenAI tool-argument limits consistently #117055
- Clear sticky ChatGPT SSE fallback when sessions end #117123
- Report provider HTTP status in diagnostic timelines #117403
- Close unread Ollama error responses cleanly #117473
- Preserve OpenAI streaming terminal delivery during Gateway shutdown #117505
- Normalize Claude tool IDs for GitHub Copilot #117789
- Preserve WebSocket transport for Codex OAuth sessions #117832
- fix(agents): close proxy SSE bodies on terminal and errors #117887
- Keep explicit llama.cpp HTTP routes on their configured transport #118293
- Return the correct final agent failure without extra delay #118331
- Report failed OpenAI-compatible chat streams as API errors #118428
- Recover Gemini streams stalled before the first response body #118511
- Keep llama.cpp auxiliary completions on native transport #119082
- Let Codex-backed agents react in current conversations #119649
- Detect truncated direct Anthropic streams without breaking proxy responses #120030
- fix(ollama): reject invalid UTF-8 in streaming NDJSON responses #120240
- Fail malformed LLM streams instead of hanging turns #120508
- Support Codex app-server 0.147.0 #120594
- fix: agent.wait can return error before the same run later ends successfully #120606
- Honor explicit OpenAI service tiers over Fast mode #120696
- Reject leaked Claude tool protocol output #120734
- fix(agents): replace stale timeout payloads precisely #122036
- Prevent replay after observed CLI timeout activity #122153
- Speed up streamed assistant output #122206
- Hide HTML error pages that follow HTTP reason phrases #122290
- Restore cached OpenAI WebSocket continuation #122483
- fix(tool-call-repair): stop re-parsing the whole response per candidate delta #122513
- fix(ai): preserve structured error code through Responses failed-event normalization #122579
- Honor the selected transport for embedded OpenAI turns #122946
- Finish agent responses after post-tool provider overloads #123005
- Retry structured transient network failures #123042
- Preserve OpenAI terminal errors and avoid pointless failover #123151
- Honor provider timeouts during stuck-session recovery #123877
- Return failures for failed OpenAI-compatible HTTP runs #124039
- Use the bundled Codex app-server for managed turns #124124
- fix(agents): continue uncached when Google prompt-cache responses fail #124479
- Preserve OpenAI WebSocket failure details and retry behavior #124591
- Keep valid Codex tools when one tool name is unsupported #124932
- Prevent duplicate audit starts for Codex dynamic tools #124976
- fix(cli-runner): drop stock watchdog defaults that disable resume promotion (#125045) #125085
- Deliver only the final answer after resumed reasoning #125149
- Propagate local-model stream policy to diagnostic recovery (#125147) #125388
- Report Ollama HTTP responses to diagnostics #125834
- Stop Ollama response handling after cancellation #125961
- Record provider request acceptance consistently #126028
- fix(agents): stop fallback tasks hanging after primary model failures #126566
- refactor(process): supervise Claude node invocations #127475
- fix(codex): accept bounded upstream prompt provenance #127730
- fix(mistral): retain provider-returned streaming model as responseModel #128066
- fix(ai): preserve tool-result boundary whitespace for Mistral and Ollama #128266
- fix(ai): keep SSE cancellation from blocking outcomes #128488
- fix(agents): reject hidden output across model verification and system turns #128570
- fix(cli): normalize echoed binary tool outputs #128699
- fix(anthropic): show Claude questions as interactive prompts #128729
- fix(agents): stop double-sanitizing openai-responses tool-call ids #128764
- fix(deepseek): doubled-bar DSML tool calls are delivered as text and never executed #128882
- fix: prevent custom Ollama cloud models from stalling indefinitely #129051
- fix(mistral): report provider failures instead of silent success #129256
- fix(ai): enforce Anthropic stream completion by provider ownership #129291
- fix(codex): back off between app-server startup retries #129505
- fix(gateway): report failed chat streams as errors #129787
- fix(ai): report interrupted local-model streams instead of empty success #129794
- fix(inbound-meta): avoid Anthropic billing classification #130298
- fix(ai): replace streamed function name on same-id continuation #130537
- fix: release stalled captured provider responses promptly #130804
- fix(copilot): keep Gemini working without live discovery #131143
- fix(copilot): own BYOK response stream lifecycles #131546
- fix(agents): show Claude CLI replies while they stream #132002
- fix: prevent first Grok tool requests from stalling #132475
- fix: prevent cold image analysis from timing out during setup #133045
- fix: prevent duplicate commentary after tool handoffs #133247
- fix(gateway): report timed-out HTTP agent runs as failures #133275
- Bound shared Codex app-server startup waits #89442
- Stop fallback models from re-running completed agent work #92011
- Make Claude CLI max-turn failures actionable and safe to retry #94130
- Retry transient provider failures after earlier tool use #97966
- Preserve complete streamed tool calls without finish_reason #98124
- Warn without interrupting Codex projection when protocols drift #99287
- Drop tool-result images for text-only Anthropic models #102329
- Keep truncated provider errors valid Unicode #102496
- Preserve emoji when Codex tool output is truncated #102522
- Keep long Mistral errors Unicode-safe #102539
- Keep Codex supervisor error messages valid around emoji #102590
- Keep Ollama malformed-stream warnings Unicode-safe #102593
- Keep truncated Codex dynamic-tool output valid Unicode #102622
- Preserve complete characters in truncated Codex tool output #104116
- Honor explicitly configured OpenAI Completions transport #105539
- Sanitize ClawRouter attribution headers safely #106454
- Preserve emoji boundaries in diagnostic timeline attributes #106986
- Use the Anthropic SDK's SSE parser #107460
- Bound oversized Ollama stream records #107473
- Preserve Unicode in Codex app-server failure messages #108335
- Restore llama.cpp compatibility for the cron tool schema #108360
- Preserve Unicode in local provider startup errors #108518
- Preserve Unicode in Codex sandbox HTTP output #108618
- Report invalid UTF-8 in provider JSON accurately #108849
- Preserve Unicode in Codex app-server error tails #109013
- Report provider finish errors accurately #109313
- Keep Unicode intact in truncated provider errors #109786
- Prevent stream cleanup errors from overriding cancellation #110427
- Keep OpenAI tool-call ID truncation Unicode-safe #110471
- Stop retrying non-retryable ChatGPT Responses errors #110655
- Reject non-text Codex WebSocket response frames #111138
- Release Ollama stream locks after early exits #111641
- Show provider 503 failures in group and channel chats #113656
- Preserve Gemini compatibility flags on discovered model variants #114831
- Preserve emoji in streamed tool arguments #115536
- Separate streamed Chat Completions content from the finish event #115718
- Preserve completed Codex answers when tool bookkeeping is incomplete #115849
- Prevent native Anthropic attempts from carrying refusal markers #116433
- Release Amazon Bedrock streaming clients #117039
- Restore
/toolsinventory for dynamically prepared models #119306 - Keep ACP retry statuses scoped to one logical dispatch #120560
- Accept TLS-certificate cloud-worker fallback events #121285
- Preserve provider local-service environment overrides on Windows #122027
- Unify CLI failover policy across execution lanes #122164
- Preserve reply cleanup when backend cancellation throws #122252
- Retry transient Undici socket failures #122535
- Recover rejected OpenAI WebSocket continuations #122727
- Accept headerless text files in Open Responses URLs #123431
- Close unhealthy ACP runtimes before replacement #126494
- fix(cli): report exact Claude stream failures #127169
- fix(gateway): Responses API returns blank or truncated assistant output #128295
- fix(ai): record provider-returned model in managed openai completions stream #128324
- fix(gateway): streamed tool calls silently omit assistant commentary #128353
- fix(gateway): retain completed status in streamed tool snapshots #128510
- fix(cli): convert HEIC sequences in model runs #128578
- chore(deps): refresh dependencies after seven-day cooldown #129187
- fix(codex): suppress under-development features warning in chat #129493
- fix(google): retain provider-returned response model #129606
- fix(ai): local-model tool calls fail in standalone library streams #129664
- fix(llama-cpp): restore tools for fallback-capable chat templates #129684
- fix: proxied model streams drop all events when the server omits the space after data: #129905
- fix(codex): keep unsupported service-tier warnings out of chat #131190
- fix(codex): hide managed Code Mode compatibility warning #131943
- fix(agents): neutralize Anthropic refusal marker replacement #132556
- fix(gateway): send finalized text with required tool calls #133206
Documentation
- Document llama.cpp schema compatibility for custom endpoints #117120
Reasoning and context limits
Reasoning settings now stay attached to the selected model and runtime when a conversation is restored or its authentication route is rebuilt. Native Codex supports Ultra for Sol and Terra and Max for Luna; embedded OpenClaw maps Ultra to the provider's highest supported effort and adds guidance for delegated work, which is a different behavior from native Codex Ultra.
Supported GPT and Claude routes expose larger context options. Normal GPT-5.5 and GPT-5.6 runs use a 272,000-token budget with an opt-in 922,000-token input window, and the Control UI can choose 200K or 1M for supported Claude 5 CLI conversations. These options remain limited to the routes, interfaces, and accounts that support them.
Compaction now uses the latest trustworthy context count rather than accumulated cache billing or duplicated history, preserves tool output across supported Responses checkpoints, and falls back to full history when a stored checkpoint is rejected. Context limits are now configured per model, and Doctor can migrate supported provider-level settings tied to explicit model entries.
Sources and complete change list
Improvements
- Keep sent prompt history stable within a session #102610
- Reduce recurring prompt and tool-instruction overhead #105095
- Make million-token OpenAI context an explicit opt-in #112916
- Consolidate model context budgets #124665
- feat(agents): selectable Claude CLI context window (200K/1M) #127951
- Add GPT-5.6 Ultra across OpenClaw and Codex runtimes #98021
- Update the managed Codex app-server to 0.144.6 #110821
- Reduce repeated per-turn conversation metadata tokens #113616
- Stabilize Codex tool prompts and expose cache losses #115238
- Add Gemini 3.7 Flash with compatible thinking levels #123366
- Add session thinking level to the bundled Swift protocol client 0d215c1
- Show the effective thinking level in /model summaries #111709
- Preserve prompt-cache reuse in history-limited sessions #115281
- Reduce duplicate ambient-room prompt content #119860
Bug fixes
- fix(agents): honor account-scoped history limits instead of silently ignoring them (#119326) 3b3372b
- Prevent stale Apple chat thinking after model switches and reconnects #103767
- Serialize chat model and reasoning changes #106534
- Stop cumulative cache usage inflating session context #108323
- Keep ACP sessions from failing on unsupported automatic thinking controls #108392
- Correct Claude 1M limits for direct Anthropic sessions #108751
- Prevent premature compaction in tool-heavy sessions #110297
- Preserve context for policy-routed CLI sessions #115284
- Enable custom providers to expose Max and Ultra reasoning #115690
- Preserve Claude CLI conversation context across prompt updates #116292
- Preserve Ollama thinking through reply maintenance #116963
- Keep runtime context focused on the active user request #117364
- Honor per-agent context limits in embedded runs #120343
- Let Claude CLI sessions compact with subscription authentication #120496
- Prevent premature compaction from inflated token counts #120497
- Preserve full Responses history across compaction retries #120729
- Align server and local compaction decisions #123397
- Prevent premature compaction from double-counted context #124267
- Let tool-heavy sessions survive runtime cutover #125181
- Apply Claude CLI thinking and preserve warm sessions #125528
- Preserve GPT-5.6 Max and Ultra selections through Codex #126492
- Use real Ollama Cloud context windows and capabilities #126653
- fix(acp): preserve configured thinking across session reuse #128169
- fix: Responses sessions compact before reaching context limit #130993
- fix: preserve ACP model switches and reply completion #132040
- fix(compaction): preserve headroom for local chat follow-up turns #132190
- Cap oversized Codex tool output before replay #99578
- Default GPT-5.6 Sol to medium reasoning e9c6c15
- Use GitHub Copilot’s actual prompt limits #103275
- Let Claude CLI choose adaptive thinking effort #103815
- Show model-specific thinking levels in Android chat #103893
- Apply provider-level token limits to configured models #105807
- Persist iOS chat thinking selections #106287
- Preserve Fable 5 effort settings on Claude CLI routes #107636
- fix(github-copilot): preserve catalog thinking efforts in requests #107834
- Use Z.AI's documented thinking payload #108633
- Honor thinking off for Kimi Code K3 #109335
- Complete Claude transcript pages across short file reads #109431
- Persist Codex automatic compaction counts #110587
- Default missing Anthropic output-token limits #110617
- Let Apple Watch chat inherit configured thinking defaults #111301
- Honor provider-specific reasoning effort mappings #112632
- Restore full verbosity and streamed reasoning controls in the TUI #113782
- Read Anthropic token limits during live model discovery #113784
- ACP turns continue when optional thinking values are rejected #113935
- Keep Copilot tool transcripts complete after interruptions #114403
- Prevent false idle timeouts during hidden reasoning #114997
- Make auto-compaction cut points free real context #115309
- Recover channel conversations from expired Claude CLI sessions #115513
- Preserve native GPT limits for Microsoft Foundry deployments #115542
- Use self-hosted providers' advertised context windows #115566
- Restore tagged Claude CLI reasoning delivery #117264
- Honor stored model overrides and cached tokens in
/compact#117580 - fix(bedrock): requests respect model output limit when reasoning is off #119155
- Stop bodyless 400 errors from triggering false compaction #119596
- fix(agent-core): preserve explicit reasoning off through Agent Core handoff #120020
- Honor Claw Router model reasoning effort metadata #120631
- Preserve long Responses sessions across cloud-worker handoff #120803
- Enable native max thinking for supported Ollama Cloud models #121074
- Preserve GPT-5 personality settings during doctor migration #121346
- Preserve Claude CLI context across heartbeats #121509
- Restore all GPT-5.6 Sol thinking levels #121865
- Resolve configured model aliases during compaction #122422
- Refresh Codex context totals after native compaction #123478
- Stabilize prompt caching across discovery order #123543
- Honor per-model context caps on Codex-routed runs #124735
- Restore GPT-5.6 reasoning effort options in Codex #126182
- fix(agents): prevent persona drift with generic slot IDs #128559
- fix(agents): native runs ignore the selected 200K context window #128838
- fix(models): preserve provider-defined reasoning levels across model catalogs #129623
- fix(agents): keep active hidden model streams out of stalled recovery #130140
- fix(ai): stabilize runtime context across Responses tool rounds #130166
- fix(ai): carry the Responses system prompt via instructions, not input[0] #130206
- fix(failover): oversized Groq requests stall the session with unusable rate-limit advice #130275
- fix(copilot): prevent static GPT-5.4 mini thinking failures #131160
- fix(anthropic): reload compacted Claude sessions #131288
- fix(agents): show token limits for rejected model requests #131990
- fix: prompt cache collapses in long sessions — aggregate tool-result truncation rewrites already-sent history #132017
- fix: keep CJK compaction summaries within budget #132145
- fix(codex): honor thinking off on supported Platform routes #132276
- fix: keep session thinking controls aligned with their model owner #132389
- fix: preserve model thinking policy and repair live validation #132431
- fix(google): preserve Gemini native compaction ownership #132527
- fix: keep summary checks after custom compactor timeouts #132536
- fix: preserve compact focus in split-turn summaries #132564
- fix(compaction): recover split-turn summaries without duplicate sections #132618
- Honor configured context limits for Claude CLI sessions #93198
- Enable maximum thinking for OpenCode Go DeepSeek V4 #99643
- Prevent silent reasoning turns from bricking Anthropic sessions #99772
- Keep Gemini thinking disabled when reasoning clamps to off #101832
- Preserve session thinking levels during unrelated updates #102866
- Match canonical reasoning efforts regardless of casing #102993
- Keep Claude fallback summaries valid at emoji boundaries #104454
- Honor Z.AI GLM-5.2 max thinking through the gateway #104778
- Enforce safe output limits for known Mistral models #105040
- fix(extra-params): preserve resolved cacheRetention against undefined own-property clobber #106069
- Show the active model and thinking level in the TUI #108669
- Preserve thinking levels in status for discovered Ollama models #108789
- Keep Claude CLI Mythos 5 context at its fallback limit #110461
- Recover from Z.AI token-limit errors #111744
- TUI footer reflects the launch thinking override #112237
- Preserve Ollama cloud model context windows #112430
- Let native iOS Talk inherit the session thinking level #112901
- Normalize standalone AI package transcripts by default #114843
- fix(agents): resolve configured context tokens for provider-self-prefixed refs #116521
- Reject invalid Chat Completions token caps #116563
- Preserve thinking levels for live-discovered Ollama models #116584
- Use Kilocode's serving-provider context limits #118868
- Preserve Chutes serving context limits #119196
- Retain images within tool-result context limits #120808
- Make
/thinksuggestions match the active model #123507 - Preserve prompt-cache boundaries for long OpenAI session IDs #126380
- fix: show context window switcher on new sessions #128621
- fix(apple): surface rejected chat session settings #128737
- fix(ai): recover input-only provider context overflows #129536
- fix(lmstudio): preserve saved model reasoning and code mode #129675
- fix(github-copilot): allow xhigh and max thinking levels on claude-opus-5 #129690
- fix(tui): make provider-specific thinking labels selectable #129952
- fix(deepinfra): preserve manifest reasoning during live discovery #129998
- fix(ollama): stop replaying images to text-only models #132106
- fix(compaction): restore fallback after safeguard provider failures #132449
- fix: keep CLI context budgets stable across provider discovery #132544
- Warn when prompt cache retention is invalid #94494
Documentation
Usage, limits, and pricing
OpenClaw now separates subscription-plan information from estimated API cost. Chat can show plan windows, reset times, credits, and the account email attached to a snapshot, while completed iOS replies can show supported input, output, cache, cost, and context-pressure details. These figures are snapshots or estimates, not provider invoices, and in mixed API-key and subscription setups the account label identifies the plan snapshot rather than every run.
The Control UI adds a Profile page for lifetime activity recorded by OpenClaw, with Usage and Profile views grouped in the selected time zone, and plugin-initiated model calls now contribute to aggregate totals. OpenClaw cannot reconstruct activity it never recorded.
Supported failed or incomplete turns can retain the provider's token and cost data without being marked successful. Permanent authentication, model, media, and long-window quota failures stop retrying, while transient rate limits and retryable server errors keep their existing retry or authorized same-provider fallback behavior; Codex subscription runs do not silently switch to pay-as-you-go API keys.
Sources and complete change list
Improvements
- Show subscription plan limits in the chat context popover #102784
- Add a lifetime activity profile to Control UI settings #102842
- Show token, cost, and context usage under iOS replies #103471
- Show the account behind plan usage #106200
- Focus Profile on identity and move analytics to Usage #115245
- Label Codex usage windows with the account email #106500
- Trace Claude Code CLI turns in model-call telemetry #108304
- Source model pricing from the hosted catalog #114060
- Speed up repeat Usage page loads with short-lived caches #115369
- Label onboarding CLI auth as subscription or usage-billed #122140
- Make cached provider usage return immediately #125520
Bug fixes
- Wait for complete all-agent usage totals #103998
- Stop retries for permanent provider errors and long quotas #105258
- Include plugin LLM calls in usage diagnostics #107937
- Make Codex quota failover model-aware and billing-safe #108254
- Recover automatically from temporary provider overloads #109354
- Harden provider caching, retries, and overflow recovery #109796
- Avoid unused Claude CLI observability overhead and bound diagnostics #110246
- Bound Control UI session-observer model calls #114052
- Restore trajectory capture for Codex sessions #115220
- Show the real Codex context window and simplify its popover #121491
- restore Codex subscription usage in
/status7eeb109 - Make provider usage load without waiting on slow providers e9620fb
- Limit all-agent session usage scans #102013
- Show periodic usage-limit errors in group chats #102643
- Keep Usage and Profile day buckets correct across DST #103097
- Correct DeepSeek V4 pricing and legacy alias metadata #103241
- Recover Groq Llama turns rejected by TPM limits #104904
- Honor Retry-After floors through one retry scheduler #105789
- Keep Usage insights aligned with active filters #105792
- Retry provider operations after HTTP 429 responses #106365
- Stop provider usage reads from hanging after headers #107454
- Report truthful Codex context freshness #107813
- Count instant sessions in hourly usage metrics #108160
- Honor the selected time zone in Usage session details #108579
- Record usage and partial output for incomplete Responses turns #109904
- Improve Ollama CJK usage estimates #110073
- Avoid false Codex subscription-exhaustion messages when account limits remain available #110381
- Honor Anthropic Retry-After cooldowns #111072
- Keep Claude usage refreshes stable on malformed responses #111088
- Report Chutes cache-read pricing correctly #111253
- Price usage with each agent's model registry #111389
- Count OpenRouter cache-write tokens correctly #111435
- Show accurate quota and rate-limit alerts #112047
- Reject malformed UTF-8 in provider usage responses #112081
- Show current MiniMax quota windows correctly #115030
- Restore CLI token usage for OpenAI-compatible providers #115564
- Correct OpenRouter budgets and Bedrock cache accounting #116896
- Correct cached-token and cost accounting for managed OpenAI chat #117043
- Correct managed Anthropic cache accounting through shared streaming policy #117086
- Report complete Responses token usage #117533
- Keep multi-call token totals cumulative #117930
- Classify exhausted xAI credits as billing failures #118615
- Apply quota suspensions to the session's real owning agent #120262
- Correct LongCat cache pricing and add provider branding #121369
- Unify provider-error handling and prevent false rate-limit matches #121817
- Correct GPT-5.5 fast-mode session costs #123152
- Restore chat cost and usage empty states #123269
- Keep provider usage tied to reloaded account config #124131
- Restore Codex harness trajectory persistence #125941
- Compact recovery ownership and model metadata #126751
- fix(announce): report unknown token usage instead of zero on the stats line #126976
- fix(usage): preserve agent ownership for matching session IDs #128803
- fix: restore usage for prepared provider credentials #129220
- fix(ui): render model provider controls progressively #130546
- fix: include compaction and branch-summary tokens in run usage #131115
- fix: reply status shows missing or stale context estimates #131554
- fix(usage): retain reset history and unify accounting #131667
- fix: usage dashboard loads slowly for large session histories #132064
- fix(ui): show plan usage only for the session provider #132972
- Keep selected-day usage totals free of undated historical records #89745
- Retry transient bare 429 provider errors #99097
- fix(agents): preserve native session accounting at finalization beb07f5
- Preserve explicit zero usage in model-call telemetry #100115
- Prevent Sonnet 5 model listings from crashing when cost data is missing #102186
- Ignore malformed provider Retry-After dates #102987
- Classify model-provider telemetry as client spans #104211
- Prevent usage-cost audits from stalling #104823
- Validate usage-footer fixed precision without breaking high precision #105297
- Keep the Usage page working during cache refreshes #105611
- Handle malformed Gemini quota responses safely #106420
- Drop invalid provider usage reset timestamps #106421
- fix(ui): usage page stays loading after cache rebuild #108270
- Cap usage history requests before Date overflow #110552
- Handle malformed Z.AI usage responses safely #110741
- Handle malformed DeepSeek balance responses #110785
- Prevent malformed Copilot usage payloads from breaking refreshes #110795
- Reject corrupted ClawRouter usage responses #111183
- Preserve throttling details in fallback notices #112702
- Reject incomplete usage date ranges #113259
- Preserve usage for failed OpenAI Responses #115175
- Normalize partial model costs before catalog publication #116317
- Preserve Claude thinking-token usage #116993
- Correct Codex app-server token usage breakdowns #117013
- fix(ios): show recent usage days first #117222
- Honor configured pricing for Claude Sonnet 5 and Opus 5 #117275
- Keep pending usage sessions visible in UTC mode #118004
- ClawRouter shows a stable error for malformed usage responses #119580
- Preserve cumulative CLI usage in turn diagnostics #120100
- Keep failover messages aligned with failure reasons #121717
- Include system-agent costs in all-agent usage totals #123271
- Keep healthy provider usage visible after auth failures #124367
- Keep provider usage tied to the refreshed account #125709
- Preserve terminal facts in oversized Codex trajectories #126050
- Preserve verified session context limits during catalog gaps #126422
- Restore
/usage costin terminal sessions #126668 - Refresh Control UI locale catalogs #126927
- fix(ui): keep Usage results aligned with active filters #127720
- fix(cli): reject malformed usage-cost --days instead of silently defaulting #127978
- fix(ui): usage filters hide sessions when multiple providers are selected #129482
- fix(kilocode): paid default router incorrectly reports zero cost #129546
- fix(ui): usage copy actions silently hide clipboard failures #129562
- fix(venice): preserve authoritative model pricing #129694
- fix(usage): prevent model identity collisions #130307
- fix(ui): Usage Shift-click selects the wrong session range #130375
- fix(usage): preserve MiniMax reset alias order #132897
Documentation