On this page

On this page

Providers

LM Studio

LM Studio runs llama.cpp (GGUF) or MLX models locally, as a GUI app or the headless llmster daemon. For install and product docs, see lmstudio.ai.

Quick start

  • Install and start the server

    Install LM Studio (desktop) or llmster (headless), then start the server:

    bash
    lms server start --port 1234

    Or run the headless daemon:

    bash
    lms daemon up

    If using the desktop app, enable JIT for smooth model loading; see the LM Studio JIT and TTL guide.

  • Set an API key if auth is enabled

    bash
    export LM_API_TOKEN="your-lm-studio-api-token"

    If LM Studio authentication is disabled, leave the API key blank during setup. See LM Studio Authentication.

  • Run onboarding

    bash
    openclaw onboard

    Choose LM Studio, then pick a model at the Default model prompt.

    The server URL prompt also accepts host shorthand such as localhost:1234. Invalid URLs stay in the prompt so you can correct them before model discovery.

    On a fresh guided setup, OpenClaw first queries /api/v1/models on the default or configured LM Studio host. An existing LLM is offered automatically only when LM Studio reports tool training and at least 16K of effective context. For loaded models, the loaded instance context takes precedence over the larger advertised maximum. The same CLI/macOS setup ladder verifies the route with a real completion before saving it. The automatic check never downloads a model and ignores embedding-only catalog entries.

  • Change the default model later:

    bash
    openclaw models set lmstudio/qwen/qwen3.5-9b

    LM Studio model keys use an author/model-name format (e.g. qwen/qwen3.5-9b); OpenClaw model refs prepend the provider: lmstudio/qwen/qwen3.5-9b. Find the exact key for a model by running the command below and looking at the key field:

    bash
    curl http://localhost:1234/api/v1/models

    After installing another model in LM Studio, refresh OpenClaw's model list:

    bash
    openclaw models list --provider lmstudio --refresh

    In the default merge mode, refresh adds discovered models while preserving your configured rows and their authored metadata, including names and context limits. An empty configured models array also supports discovery on an unauthenticated server. models.mode: "replace" keeps only explicitly configured models.

    Non-interactive onboarding

    bash
    openclaw onboard --non-interactive --accept-risk --skip-health --auth-choice lmstudio

    Or specify base URL, model, and API key explicitly:

    bash
    openclaw onboard \  --non-interactive \  --accept-risk \  --skip-health \  --auth-choice lmstudio \  --custom-base-url http://localhost:1234/v1 \  --lmstudio-api-key "$LM_API_TOKEN" \  --custom-model-id qwen/qwen3.5-9b

    --custom-model-id takes the model key as returned by LM Studio (e.g. qwen/qwen3.5-9b), without the lmstudio/ provider prefix. Pass --lmstudio-api-key (or set LM_API_TOKEN) for authenticated servers; omit it for unauthenticated servers and OpenClaw stores a local non-secret marker instead. --custom-api-key is still accepted for compatibility, but --lmstudio-api-key is preferred.

    This writes models.providers.lmstudio and sets the default model to lmstudio/<custom-model-id>. Providing an API key also writes the lmstudio:default auth profile.

    Add --json for a machine-readable result. Connection, HTTP, and model-selection failures return a JSON error with the same recovery guidance as human output and exit nonzero without applying the proposed provider configuration.

    Interactive setup can additionally prompt for a preferred load context length and applies it across the discovered models it saves to config.

    Configuration

    Streaming usage compatibility

    LM Studio doesn't always emit an OpenAI-shaped usage object on streamed responses. OpenClaw recovers token counts from llama.cpp-style timings.prompt_n / timings.predicted_n metadata instead. Any OpenAI-compatible endpoint resolved as a local endpoint (loopback host) gets this same fallback, which covers other local backends such as vLLM, SGLang, llama.cpp, LocalAI, Jan, TabbyAPI, and text-generation-webui.

    Gemma 4 tool-call recovery

    For Gemma 4 models using openai-completions, OpenClaw recovers complete standalone <|tool_call>call:...<tool_call|> batches if the server returns them as text, including when the call is the last content before finish_reason: "stop". Recovery preserves raw string arguments and requires complete argument objects; incomplete calls, prose, and code examples remain text. Truncated, filtered, cancelled, or unterminated streams do not authorize recovered calls. If native tool calls also appear in the stream, they remain authoritative and raw text is not promoted.

    Thinking compatibility

    When LM Studio's /api/v1/models discovery reports model-specific reasoning options, OpenClaw exposes matching reasoning_effort values (none, minimal, low, medium, high, xhigh) in model compat metadata. Some LM Studio builds advertise a binary UI option (allowed_options: ["off", "on"]) while rejecting those literal values on /v1/chat/completions; OpenClaw normalizes that binary shape to the six-level scale before sending requests, including for older saved config that still has off/on reasoning maps.

    For graded options, fresh discovery maps max (and Ultra's provider effort) to the highest advertised canonical effort, regardless of option order. For example, off, low, medium maps max to medium. Existing saved graded reasoningEffortMap values remain explicit configuration; rerun LM Studio setup to regenerate them from current server metadata.

    Explicit configuration

    json5
    {  models: {    providers: {      lmstudio: {        baseUrl: "http://localhost:1234/v1",        apiKey: "${LM_API_TOKEN}",        api: "openai-completions",        models: [          {            id: "qwen/qwen3-coder-next",            name: "Qwen 3 Coder Next",            reasoning: false,            input: ["text"],            cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },            contextWindow: 128000,            maxTokens: 8192,          },        ],      },    },  },}

    Model instances and context

    With preload enabled, OpenClaw routes chat requests to a loaded instance with enough context for the selected model budget. A newly loaded instance is addressed by the identifier returned by LM Studio. Your configured model reference and conversation model identity keep the canonical model key.

    Model loads use the configured provider timeoutSeconds (or the request timeout override), with a two-minute default matching embedding loads. Increase models.providers.lmstudio.timeoutSeconds for slow cold loads. If a load fails while every known loaded instance is too small, OpenClaw reports the model and requested context instead of sending the prompt to a smaller instance. Wait for loading to finish in LM Studio and retry, or lower the model's contextTokens.

    With preload enabled, embedding requests also check that their model is loaded and route to the instance prepared for the configured context length. This avoids truncating input through a smaller loaded instance and lets memory embeddings recover after model eviction even when LM Studio JIT loading is disabled. Embedding model and cache identity keep the canonical model key.

    Disabling preload

    LM Studio supports just-in-time (JIT) model loading, loading models on first request. OpenClaw preloads models through LM Studio's native load endpoint by default, which helps when JIT is disabled. To let LM Studio's JIT, idle TTL, and auto-evict behavior own model lifecycle instead, disable OpenClaw's preload step:

    json5
    {  models: {    providers: {      lmstudio: {        baseUrl: "http://localhost:1234/v1",        api: "openai-completions",        params: { preload: false },        models: [{ id: "qwen/qwen3.5-9b" }],      },    },  },}

    LAN or tailnet host

    Use the LM Studio host's reachable address, keep /v1, and make sure LM Studio is bound beyond loopback on that machine:

    json5
    {  models: {    providers: {      lmstudio: {        baseUrl: "http://gpu-box.local:1234/v1",        apiKey: "lmstudio",        api: "openai-completions",        models: [{ id: "qwen/qwen3.5-9b" }],      },    },  },}

    lmstudio automatically trusts its configured endpoint for model requests, including loopback, LAN, and tailnet hosts (except metadata, link-local, and local-use NAT64 64:ff9b:1::/48 origins). Any custom/local OpenAI-compatible provider entry gets the same exact-origin trust. Requests to a different private host or port still require models.providers.<id>.request.allowPrivateNetwork: true; set it to false to opt out of the default trust.

    Troubleshooting

    Model discovery failures

    When a configured server cannot list models, OpenClaw reports an unavailable catalog or a catalog authentication rejection. A refresh can keep the last successful inventory when the connection and credentials still match. A successful empty response clears discovered models; explicitly configured models remain available without discovery. Restore the server connection or correct its credentials, then refresh the model list.

    LM Studio not detected

    Make sure LM Studio is running:

    bash
    lms server start --port 1234

    If authentication is enabled, also set LM_API_TOKEN. Verify the API is reachable:

    bash
    curl http://localhost:1234/api/v1/models

    Authentication errors (HTTP 401)

    • Check that LM_API_TOKEN matches the key configured in LM Studio.
    • See LM Studio Authentication.
    • If the server does not require authentication, leave the key blank during setup.
    Was this useful?