On this page
On this page
Providers
LM Studio
LM Studio runs llama.cpp (GGUF) or MLX models locally, as a GUI app or the headless llmster
daemon. For install and product docs, see lmstudio.ai.
Quick start
Install and start the server
Install LM Studio (desktop) or llmster (headless), then start the server:
Or run the headless daemon:
If using the desktop app, enable JIT for smooth model loading; see the LM Studio JIT and TTL guide.
Set an API key if auth is enabled
If LM Studio authentication is disabled, leave the API key blank during setup. See LM Studio Authentication.
Run onboarding
Choose LM Studio, then pick a model at the Default model prompt.
The server URL prompt also accepts host shorthand such as localhost:1234.
Invalid URLs stay in the prompt so you can correct them before model discovery.
On a fresh guided setup, OpenClaw first queries /api/v1/models on the
default or configured LM Studio host. An existing LLM is offered automatically
only when LM Studio reports tool training and at least 16K of effective
context. For loaded models, the loaded instance context takes precedence over
the larger advertised maximum. The same CLI/macOS setup ladder verifies the
route with a real completion before saving it. The automatic check never
downloads a model and ignores embedding-only catalog entries.
Change the default model later:
LM Studio model keys use an author/model-name format (e.g. qwen/qwen3.5-9b); OpenClaw model refs
prepend the provider: lmstudio/qwen/qwen3.5-9b. Find the exact key for a model by running the
command below and looking at the key field:
After installing another model in LM Studio, refresh OpenClaw's model list:
In the default merge mode, refresh adds discovered models while preserving your
configured rows and their authored metadata, including names and context limits.
An empty configured models array also supports discovery on an unauthenticated
server. models.mode: "replace" keeps only explicitly configured models.
Non-interactive onboarding
Or specify base URL, model, and API key explicitly:
--custom-model-id takes the model key as returned by LM Studio (e.g. qwen/qwen3.5-9b), without
the lmstudio/ provider prefix. Pass --lmstudio-api-key (or set LM_API_TOKEN) for authenticated
servers; omit it for unauthenticated servers and OpenClaw stores a local non-secret marker instead.
--custom-api-key is still accepted for compatibility, but --lmstudio-api-key is preferred.
This writes models.providers.lmstudio and sets the default model to lmstudio/<custom-model-id>.
Providing an API key also writes the lmstudio:default auth profile.
Add --json for a machine-readable result. Connection, HTTP, and model-selection
failures return a JSON error with the same recovery guidance as human output and
exit nonzero without applying the proposed provider configuration.
Interactive setup can additionally prompt for a preferred load context length and applies it across the discovered models it saves to config.
Configuration
Streaming usage compatibility
LM Studio doesn't always emit an OpenAI-shaped usage object on streamed responses. OpenClaw
recovers token counts from llama.cpp-style timings.prompt_n / timings.predicted_n metadata
instead. Any OpenAI-compatible endpoint resolved as a local endpoint (loopback host) gets this same
fallback, which covers other local backends such as vLLM, SGLang, llama.cpp, LocalAI, Jan, TabbyAPI,
and text-generation-webui.
Gemma 4 tool-call recovery
For Gemma 4 models using openai-completions, OpenClaw recovers complete standalone
<|tool_call>call:...<tool_call|> batches if the server returns them as text,
including when the call is the last content before finish_reason: "stop".
Recovery preserves raw string arguments and requires complete argument objects;
incomplete calls, prose, and code examples remain text. Truncated, filtered,
cancelled, or unterminated streams do not authorize recovered calls. If native
tool calls also appear in the stream, they remain authoritative and raw text is
not promoted.
Thinking compatibility
When LM Studio's /api/v1/models discovery reports model-specific reasoning options, OpenClaw
exposes matching reasoning_effort values (none, minimal, low, medium, high, xhigh) in
model compat metadata. Some LM Studio builds advertise a binary UI option (allowed_options: ["off", "on"]) while rejecting those literal values on /v1/chat/completions; OpenClaw normalizes that
binary shape to the six-level scale before sending requests, including for older saved config that
still has off/on reasoning maps.
For graded options, fresh discovery maps max (and Ultra's provider effort) to the highest
advertised canonical effort, regardless of option order. For example, off, low, medium
maps max to medium. Existing saved graded reasoningEffortMap values remain explicit
configuration; rerun LM Studio setup to regenerate them from current server metadata.
Explicit configuration
Model instances and context
With preload enabled, OpenClaw routes chat requests to a loaded instance with enough context for the selected model budget. A newly loaded instance is addressed by the identifier returned by LM Studio. Your configured model reference and conversation model identity keep the canonical model key.
Model loads use the configured provider timeoutSeconds (or the request timeout override),
with a two-minute default matching embedding loads. Increase models.providers.lmstudio.timeoutSeconds
for slow cold loads. If a load fails while every known loaded instance is too small, OpenClaw
reports the model and requested context instead of sending the prompt to a smaller instance.
Wait for loading to finish in LM Studio and retry, or lower the model's contextTokens.
With preload enabled, embedding requests also check that their model is loaded and route to the instance prepared for the configured context length. This avoids truncating input through a smaller loaded instance and lets memory embeddings recover after model eviction even when LM Studio JIT loading is disabled. Embedding model and cache identity keep the canonical model key.
Disabling preload
LM Studio supports just-in-time (JIT) model loading, loading models on first request. OpenClaw preloads models through LM Studio's native load endpoint by default, which helps when JIT is disabled. To let LM Studio's JIT, idle TTL, and auto-evict behavior own model lifecycle instead, disable OpenClaw's preload step:
LAN or tailnet host
Use the LM Studio host's reachable address, keep /v1, and make sure LM Studio is bound beyond
loopback on that machine:
lmstudio automatically trusts its configured endpoint for model requests, including loopback,
LAN, and tailnet hosts (except metadata, link-local, and local-use NAT64
64:ff9b:1::/48 origins). Any custom/local OpenAI-compatible
provider entry gets the same exact-origin trust. Requests to a different private host or port still
require models.providers.<id>.request.allowPrivateNetwork: true; set it to false to opt out of
the default trust.
Troubleshooting
Model discovery failures
When a configured server cannot list models, OpenClaw reports an unavailable catalog or a catalog authentication rejection. A refresh can keep the last successful inventory when the connection and credentials still match. A successful empty response clears discovered models; explicitly configured models remain available without discovery. Restore the server connection or correct its credentials, then refresh the model list.
LM Studio not detected
Make sure LM Studio is running:
If authentication is enabled, also set LM_API_TOKEN. Verify the API is reachable:
Authentication errors (HTTP 401)
- Check that
LM_API_TOKENmatches the key configured in LM Studio. - See LM Studio Authentication.
- If the server does not require authentication, leave the key blank during setup.