On this page

On this page

Gateway

Restart recovery

Conversations, transcripts, scheduled jobs, native subagent records, and queued outbound messages live on disk. After a gateway restart, eligible work interrupted mid-turn is detected and resumed automatically. Recovery is always on and normally needs no manual intervention. Exhausted infrastructure retries, or a missing durable message-action authority claim, may tombstone one session until you inspect or replace it.

This page describes what survives a restart, how interrupted work is detected, and what the automatic resume looks like.

What survives a restart

State Storage Behavior across restart
Conversation history Per-agent SQLite database Untouched; sessions continue from the stored transcript
Accepted Control UI follow-ups Per-agent SQLite pending inputs and browser outbox Matching interrupted inputs are re-admitted when the browser reconnects
Interrupted main-session turn Per-agent SQLite session row and transcript Automatically resumed or reconciled a few seconds after startup
Subagent runs SQLite (shared state database) Interrupted runs settle; the parent decides how to continue
Queued outbound deliveries SQLite delivery queue Drained after restart; undelivered replies are retried
Scheduled (cron) jobs SQLite cron store Schedules persist; the scheduler re-arms on boot
Restart continuation SQLite restart sentinel One-shot follow-up dispatched to the session that asked for the restart
Gateway terminal PTYs Process memory End with the old process; terminal sessions are not recovered

The Control UI retains accepted text and attachments in its outbox until the Gateway confirms transcript consumption. After reconnecting, it checks the saved receipt before resubmitting an interrupted input through normal authentication and session admission. The old queue and execution authority are never reused. This preserves each browser outbox's order without submitting already-consumed messages again. Different browsers can reconnect in a different order.

Accepted Control UI input is committed to its admitted source session before ACP execution or question consumption. If the conversation routes to a bound ACP session, the source keeps the original request and the bound session owns the ACP transcript and reply. An output-persistence failure does not make consumed input eligible for automatic resubmission.

If the browser no longer has the matching outbox payload, the saved interrupted input remains available for explicit resend. Cancelled input is not replayed. Inputs accepted by older versions without resumable custody also require explicit resend. An uncertain submission stays unconfirmed until its outcome can be reconciled or the user chooses to retry it. Recovery of a turn already in the transcript does not depend on the browser returning.

When an update replaces the bundled Control UI, an open tab reloads after the Gateway reports the new build. Automatic recovery for that reported build and manual reloads share a bounded document-readiness check, so a transient failed check does not immediately strand the tab. Generic lazy-chunk failures make one automatic check and leave further recovery to the visible retry action. The browser still limits automatic navigation to one reload per target build. If the Gateway remains unavailable, use the visible reload action once it is reachable.

Downgrading to v2026.9.2 preserves newer accepted-input records, but that version rejects same-ID retries against retained newer receipts before execution. This includes queued or interrupted inputs and consumed collected-input receipts. Consumed receipts remain excluded from pending counts. Individually consumed inputs have already left the pending-input store and retain normal transcript idempotency. Upgrading again restores matching unconsumed-input recovery through fresh authentication and admission, provided the session and accepted input have not been changed or removed. Already-consumed input remains consumed. Do not delete receipts or change message IDs merely to bypass a downgrade conflict.

Native Codex recovery reconstructs saved document contents under the current attachment policy and context limits. The Codex plugin owns that native input path: update its artifact alongside the Gateway. An intentionally pinned older plugin does not acquire the fix from a core-only update.

A detached harness completion that has entered a parent turn with an outbound channel route can retain an exact source completion, terminal outcome, and requester-session claim. After a crash, main-session recovery continues that admitted turn without starting the child again. The original completion identity remains distinct from the recovery run. A retained final receipt settles the admitted completion even if the old native monitor no longer exists; a pending recovery claim is not a delivered result. A later outcome cannot authorize the earlier completion input or settle its receipt. Cancellation, session replacement, and missing or contradictory source identities do not authorize replay. A held admission or still-valid admitted input keeps recovery pending. Native parent rotation must preserve the current connection and requester identity checks. A later parent registration cannot supply missing historical ownership.

Before parent completion admission, native Codex pending assignments retain their run, child-thread, native-parent, and known native-turn identities in the existing parent binding metadata. Accepted follow-ups can also retain their submission receipt there. No new table or Tasks record is created, and historical Tasks rows are not imported. When the parent registers again, recovery reads the saved assignments and native history under fresh host-issued completion authority. It checks the original requester session, lifecycle, connection, and child lineage before restoring observation or routing a result. Native-parent thread rotation can preserve that metadata within the same requester lifecycle and connection; a reset, requester-generation replacement, or connection-policy change does not authorize adoption. Missing historical ownership stays unknown. This recovery requires a retained assignment and a new parent registration; native lineage alone does not reconstruct an unrecorded completion obligation.

The admitted-turn recovery described above applies to newly recorded channel-delivery claims, not route-less, transcript-only, or Control UI parents. It does not recover every older native completion. Progress messages, empty or truncated receipts, and uncertain sends cannot establish final delivery. The Gateway and Codex plugin must both be updated for recovery joining. Older builds may discard the added receipt fields while rewriting session metadata, even without a SQL schema change. Finish pending completion recovery before downgrading; do not delete consumed inputs or change source IDs to force another delivery.

Pending delivery rows drain or retry after restart. When a delivery exhausts its retry budget, recovery reclaims expired producer custody. An active producer keeps ownership. Failed deliveries cannot send again, but retain the information needed to settle their owning session or conversation. If that update fails or the gateway crashes, recovery resumes the update without resending the message. After settlement, failed rows discard their payload. Only reusable or crash-ambiguous owners keep a minimal bounded or permanent receipt that prevents duplicate delivery. Delivery uncertainty notices retain their acknowledgment, so a repeated settlement cannot notify the same intent again.

Finish pending settlements before downgrading. Older builds may discard their metadata during database repair or drop acknowledged notices while rewriting session records, even when the schema version is unchanged. See Database schemas for downgrade precautions.

Graceful restarts drain first

Startup migration warnings do not prevent the Gateway from starting. It logs the warnings once and starts degraded. openclaw status and openclaw doctor show the running Gateway's warning report. Read-only operators receive the repair hint. Warning details are restricted to administrators and startup logs. Run openclaw doctor --fix against the same state/config, then restart the Gateway. Unfinished migrations remain pending for a later startup. Errors that leave required state unsafe to read still stop startup.

A requested restart (openclaw gateway restart, a config change that requires a restart, or a gateway update) does not kill in-flight work immediately. The gateway stops accepting new work, then waits for active agent turns and background tasks to finish, up to a drain budget (5 minutes by default). Most restarts therefore interrupt nothing at all.

CLI shutdown drains process-wide work. Closing an individual Gateway drains its own active chat runs and queued turns, and waits on the process-wide pending-reply count. Both report remaining work as named counts; categories can overlap and should not be added as distinct turns.

When Gateway connection shutdown begins, session observer event subscriptions stop before the observer closes. Accepted observer work still settles, while chat events and terminal session writes continue through their own drain.

If database closure refuses a follow-up drain, the drain parks its queued input instead of retrying against the closing owner. A fresh drain request can resume it after database admission reopens; durable input recovery owns restart replay.

Read-only RPC waits (agent.wait, approval decision waits, question.waitAnswer, and device.scopes.waitUpgrade) stop observing when their client disconnects. When shutdown drain begins, connected waiters receive retryable UNAVAILABLE errors with reason gateway-restarting, so clients can reconnect and wait again. These waits do not consume the stop drain budget. The underlying runs, decisions, and admitted writes keep their normal drain and recovery behavior. This also applies to update restarts once the running Gateway contains this fix; installing new files cannot change a wait already held by an older Gateway process.

For an accepted WebChat turn, agent.wait reports pending with timeoutPhase: "queue" when queue admission completes, including when the wait began before admission. Queuing does not complete the turn's run ID.

Cron shutdown gives execution cleanup and durable result writes the same cleanup window. Finishing the job's execution does not by itself complete the drain: result persistence must also settle, or the Gateway reports the remaining work when the window expires. Exit-watcher failures are reported after these drains settle.

On Linux and macOS, this also applies when startup recovers from an unsupported Node version and the service manager tracks a launcher parent. The launcher forwards the stop signal and waits for the serving Gateway to drain within the shared service budget. Managed restart intent targets the live serving owner, so unfinished work still follows restart recovery when its drain budget expires. When launchd drives the stop, macOS uses the running job's own ExitTimeOut, capped by the launcher's own stop timer when a launcher is in the path, and Linux units use their own stop timeout; both are described in the deadline sections below. This requires a Gateway started with the updated launcher: replacing files cannot change a launcher that is already running. On macOS, a marked launcher parent keeps a conservative 19-second child deadline because its existing markers do not identify whether it loaded the old or new stop policy.

For these managed restarts, if the CLI cannot verify the service command, serving owner, or restart-intent recording, it refuses the restart before signaling with GATEWAY_RESTART_PREPARATION_REFUSED. Restore service inspection or state access, verify Gateway status, and retry. An unverified launcher PID is never a fallback.

Running 2026.9.4 Gateways remain eligible through their published state-local lock identity. The CLI verifies the service installation, live process start identity, and membership in that native service before preparing the restart. Stale or mismatched identities cannot authorize a running-service restart. An inactive service can still start without restart intent once the prior owner is proven dead.

On Linux, the systemd unit must use KillMode=mixed so the initial stop signal reaches only the Gateway. Systemd still kills remaining child processes when the Gateway exits or its stop deadline expires. Older KillMode=control-group units signal child runtimes immediately, which can interrupt a turn before drain finishes. The spawn broker stays available while its Gateway connection is alive, even if it receives the stop signal too, so cleanup can still launch commands and observe child exits. This does not protect other child runtimes; KillMode=mixed remains required. Updates and openclaw doctor --fix refresh outdated OpenClaw-managed Linux unit policy. Maintenance reads the resident shutdown budget from Gateway status. Older Gateways without that fact follow the short-budget path: fence admission and observe lifecycle drain until idle or the update step deadline. At the deadline, outstanding write custody refuses the stop with its owner phase; remaining turns can be interrupted with a recorded warning. Operator-owned drop-ins must be inspected and updated separately because reinstalling the base unit preserves them. See Linux services.

Maintenance custody observations

Gateway status reports its process-owned shutdownBudget, with activeWork counts and a separate writeCustody array. Suspension preparation and status responses also include optional writeCustody entries with phase and count. Current phases identify migration, backup, coordinator writes, session lifecycle mutation, and terminal persistence. These are recorded by their operation owners; ordinary root requests and cron runs do not imply write custody. Counts can overlap.

The optional field is additive. Older Gateways, including published 2026.9.5, can omit it. Missing custody information never refuses maintenance. If the update step deadline expires, maintenance stops with a warning that includes the latest root-request and cron-run counts, explains the resident's missing distinction, and identifies the next Gateway's refreshed stop policy. Only a reported live write-custody phase refuses that deadline stop.

Systemd stop deadlines

At startup and when accepting shutdown, the Gateway reads its running systemd unit's effective TimeoutStopUSec, including drop-ins. It logs the source and reconciled stop budget at both points, so a repaired unit takes effect without restarting first. Inspection and any wait for startup to finish consume the same shutdown deadline. Active-work drain uses at most 315 seconds, with up to 10 seconds reserved for final chat writes and server cleanup and up to another 5 seconds before systemd's deadline. A unit with the default 90-second stop timeout therefore gets a 75-second drain and an 85-second Gateway shutdown deadline. A shorter supervisor timeout also caps requested restart waits.

Both allowances are bounded when the deadline cannot fund them, by the same rule the launchd budget uses and for the same reason: subtracting two values sized for a 315-second drain from a short stop timeout consumed it entirely and left active work nothing. A unit whose TimeoutStopSec is 20 seconds or longer keeps the allocation it already had, less whatever the stop inspection itself cost, rather than to the exact millisecond: 20 seconds is only the least that funds the full 10-second reserve alongside a 5-second drain before that cost comes off. Below that the deadline cannot fund both, and the reserve yields to keep a drain: a unit at or below 15 seconds now drains at all where it previously drained for zero milliseconds, and the four values between trade part of a reserve for a drain that was under 5 seconds. The reserve never falls below half the shutdown budget. The exit margin holds its full 5 seconds down to a 20-second deadline and shrinks below that, to 3.75 seconds at 15 seconds and a quarter of anything shorter. The drained work, ordering, and interruption behavior stay the same, and systemd's own 90-second default is unaffected.

Service-child cleanup uses the remaining Gateway shutdown budget, leaving time for final exit bookkeeping. A forced restart drains admitted work within the same budget. Expiring the restart scheduler's idle deferral starts the normal drain: new work is fenced and admitted work gets the computed grace period, still capped by the native shutdown deadline. Deferral time does not spend that shutdown budget. For a supervisor's SIGTERM restart, a shorter requested drain limits when active runs are interrupted, not when database cleanup must finish: cleanup can use the remaining native stop budget. After an interrupted external restart, the Gateway joins final chat persistence, drains Memory's database borrows, and closes agent databases before exiting. It skips plugin and channel teardown that could retain the process until the deadline. Database admission stays fenced through lock release and log flushing, which have a final five-second exit window. Shutdown logs report each cleanup step that takes at least one second, including its name and elapsed milliseconds. Deadline exits list the currently pending cleanup steps and how long each has been waiting (up to eight, with a count of any others). These diagnostics are always enabled; OPENCLAW_GATEWAY_RESTART_TRACE=1 adds begin/end timing for fast steps too. The supervisor deadline remains the hard upper bound. Clean database restart proof is published only after writer leases, checkpointing, and native connection closure settle; one database's idle receipt alone does not authorize process exit. A restart without a supervisor handoff uses the existing shutdown deadline for cleanup. This includes foreground Gateways inside another service's cgroup, restarts with OPENCLAW_NO_RESPAWN=1, and standalone updates that must launch their own replacement. Cgroup membership alone does not provide a supervisor that will replace the Gateway. Ordinary cancellation keeps its five-second grace before forced termination. During shutdown, a relay that needs forced termination after its owned processes are confirmed gone produces a warning. Completed cleanup leaves the Gateway's exit status at zero; an unconfirmed process cleanup boundary still reports failure.

The process's cgroup selects the system or user manager, independently of the account running the Gateway or its restart owner. This also covers hand-written system units with User=openclaw and externally managed deployments. Reading the system unit's timeout does not require sudo or notification support.

If the unit cannot be inspected at startup, the Gateway warns with the manager, unit, and failure reason and uses systemd's 90-second default as a conservative fallback. A failed shutdown reread retains the startup budget, with elapsed time deducted, instead of assuming a longer timeout. An explicitly unlimited timeout keeps the normal Gateway budget. Already-running v2026.9.5 Gateways retain their startup reading until they restart; installing newer files cannot change the shutdown budget captured by that older process.

An already-installed old unit benefits from the clamp as soon as the new Gateway starts, without a service rewrite. This leaves time for orderly shutdown instead of spending the entire stop window in drain. Work that cannot settle still uses the existing interruption and recovery path; the shorter budget cannot guarantee that arbitrary cleanup completes. CLI installs and guided Doctor service repairs render TimeoutStopSec=330 from the same policy as the Gateway. Doctor reports an effective stop timeout below that requirement. Operator-owned drop-ins remain the operator's responsibility.

For hand-written system units, allow at least drain + 15 seconds. With the current maximum drain, create /etc/systemd/system/openclaw-gateway.service.d/stop-timeout.conf:

ini
[Service]TimeoutStopSec=330

Run sudo systemctl daemon-reload and verify with systemctl show openclaw-gateway.service -p TimeoutStopUSec. Restart through your service's deployment owner. For a user unit, use systemctl --user edit openclaw-gateway.service and the corresponding --user reload/show commands. Retain KillMode=mixed as described above; a longer timeout does not protect children from KillMode=control-group's initial signal.

Replies to pending node commands remain accepted during the drain, including worker cleanup started by shutdown. Each reply must still match its live invocation, node connection, pairing generation, and owning lifecycle. This lets cleanup finish without waiting for a command timeout. It does not reopen admission for new requests.

Operators can also inspect and answer pending questions or resolve approvals while the Gateway drains. These requests must belong to still-pending work admitted before shutdown; normal authorization checks still apply. New question and approval requests remain fenced.

Only work that cannot finish inside the drain budget (or any run interrupted by a forced restart or a crash) is aborted — and before that happens, each affected session is marked for recovery.

Restart cancellation also preserves recoverability when the bulk shutdown marker cannot be written. Native runtime preparation interrupted by the same Gateway restart is recorded as restart cancellation rather than a provider failure. Explicit user cancellation and genuine execution timeouts remain terminal. Recovery startup uses the admitted run's existing deadline, including runtime preparation and waiting for session or global capacity. Waiting in a healthy queue does not consume separate failed-start attempts.

sessions.abort waits for the cancellation's session write before acknowledging success. Restarting immediately after that acknowledgment preserves the terminal outcome even if the run's finalizer has not finished. This also applies to a parent that yielded while waiting for spawned tasks: successfully stopping its children records the captured parent's cancellation before acknowledging, without overwriting a newer turn in that session. If another child cannot be stopped, the response still reports incomplete cancellation; the captured parent's cancellation is persisted before that error.

Launchd stop deadlines

On a launchd-driven stop, the macOS Gateway reads the loaded job's effective exit timeout with launchctl print. It checks the system, GUI, and user domains, and accepts only a job whose PID is the Gateway or its launcher. Restart ownership can still be external: the supervisor that enforces the stop, not the owner of the next start, determines the deadline. A direct SIGTERM does not start launchd's stop clock, so it keeps the existing Gateway stop policy unless an OpenClaw launcher independently enforces its own child reap timer.

The stop budget is time available until the supervisor can kill the process. The Gateway sets aside an exit margin (up to 5 seconds) and plans its shutdown inside the remainder. Drain is the first part of that shutdown: stop admitting new work and wait for active turns and background tasks to settle. The rest is reserved for final writes and service cleanup (up to 10 seconds). The Gateway logs the chosen source, drain, shutdown deadline, reserve, and exit margin.

New LaunchAgent plists request ExitTimeOut=330, derived from the same service stop budget as Linux's TimeoutStopSec=330. With that effective deadline, the nominal split is 315 seconds of drain, 10 seconds for cleanup, and 5 seconds before the supervisor deadline. openclaw doctor --fix backs up and migrates the retired installer value of 20 seconds to 330 only when the definition otherwise matches the retired generated template. A 20-second timeout in a customized definition is ambiguous: Doctor reports that it is below the drain budget and preserves it. An installer comment alone does not authorize replacing an operator's policy. Other explicit values, including 600 and unlimited (0), are reported and preserved; short custom values receive a warning. Update finalization uses the same backed-up service reconciliation. The first stop still obeys the previously loaded job's deadline until the replacement is bootstrapped.

Time spent inspecting the job is debited. A legacy 20-second job nominally leaves 5 seconds of drain, 10 seconds for cleanup, and a 5-second exit margin. For a shorter custom deadline the exit margin is at most a quarter, and the cleanup reserve gives way to leave active work up to 5 seconds of drain (at most half of the remaining shutdown budget when it is very short). For example, a 5-second job nominally leaves 1.875 seconds each for drain and cleanup after its 1.25-second exit margin; a 15-second job leaves 5 seconds of drain, 6.25 seconds for cleanup, and a 3.75-second margin. This reduces cleanup time for custom jobs below 20 seconds. The default systemd 90-second deadline is unaffected. The check can consume up to three 2-second calls; a very slow inspection can leave no drain.

If a Node-recovery or compile-cache launcher is the job's PID, its own child reap timer may be shorter than launchd's deadline. The Gateway caps its budget at that timer only when its respawn markers establish that this launcher started it; an unrelated parent does not shorten the budget. New launchers use the shared 330-second service budget, but the child conservatively retains the older 19-second launcher cap until a parent runtime-generation contract can distinguish already-running launchers. Raising the plist value alone does not remove that cap.

A direct signal to a marked launcher can start its own timer while launchd still reports running. The Gateway cannot distinguish that forwarded signal from a direct signal to its child, where the launcher has no timer, so it does not cap this case. Operators stopping a launcher directly should use the loaded job stop instead to get an enforceable, reported deadline.

ExitTimeOut=0 means no launchd stop deadline, not a missing value. Without a launcher timer, the Gateway keeps its ordinary policy and does not classify the job as a native deadline. An unreadable job leaves the existing stop policy in place with a warning, while a confirmed stopping job whose timeout is missing uses launchd's 20-second default with a warning.

The value comes from the job launchd has loaded, not the plist on disk. Editing ExitTimeOut or running launchctl kickstart -k does not reload it; launchctl bootout followed by launchctl bootstrap does, restarting the Gateway. Reading per stop avoids a stale startup snapshot, not a stale loaded job. On macOS 27, launchd reports at most 60 seconds even if the plist asks for more; inspect the loaded job to confirm the effective value.

Host sleep and process freezes

When a gateway host wakes from sleep, a virtual machine resumes, or the process continues after a long pause, the gateway detects the freeze within about 30 seconds. It restarts channel connections once tracked Gateway work is idle, then refreshes cached health and presence. The health and presence refresh still runs when a busy gateway defers only the channel restart. This keeps stale sockets from waiting for their normal expiry without interrupting an active reply or agent startup when a busy event loop caused the timer gap.

The macOS app and Linux companion cooperate with a local gateway by preparing a short suspension lease before the host sleeps and resuming it after wake. Remote gateways are not suspended when the app host sleeps. A deliberate suspension through gateway.suspend.* keeps recovery deferred until the controller resumes the gateway.

Recovery after a failed update

After a failed interactive update or repair, OpenClaw finishes cleanup and any service recovery, then opens openclaw triage. Triage immediately starts the first directly launchable coding agent in this order: Claude Code, Codex, OpenCode, then Pi. It passes the captured failure before fresh Doctor checks or archive collection and asks the agent to diagnose, repair, and verify the installation. The agent receives the captured installation paths and keeps its normal authentication, sandbox, and approval settings.

For a failed Control UI or unattended update, use the installation-specific command printed on the Gateway host, or run triage there with the same OpenClaw profile and state/config paths. Use --agent to select a particular coding agent:

bash
openclaw triageopenclaw triage --agent codex

JSON, --yes, and non-interactive update invocations collect diagnostics without starting an external coding agent. openclaw triage --non-interactive also prepares diagnostics without launching an agent. --update-result <path> includes an updater's saved failure artifact. Printed handoff commands preserve installation selectors and use PowerShell on Windows or POSIX shells on macOS, Linux, and WSL.

Staging and validation run while the old Gateway serves. The candidate runs Doctor lint, config and plugin planning, and an isolated canary boot against copied configuration and verified database snapshots. Migrations on these copies rehearse the upgrade without changing live state. A validation failure can enter bounded unattended repair while the old Gateway keeps serving. Activation requires the failed check to pass. Otherwise the candidate is discarded. An already-current no-op never stops the Gateway. Older targets that predate migration continuation record runtime validation as unavailable and use the existing downgrade finalization path. The detached helper also waits for the activating phase before parking its parent Gateway. The first activation window contains the swap, required live migrations, and service start. Plugin package download and sync run while the core Gateway serves. A changed plugin snapshot requires a second measured activation window: full Doctor migrations under exclusive maintenance, then restart and verification. Unchanged plugins do not run another full Doctor pass.

After activation, the updater verifies that the managed service is running and owns its port, the Gateway hello handshake matches the expected version/build identity, a 12-check health settle passes, plugins and channels are healthy, and /readyz returns HTTP 200. Update verification does not use model inference. Startup receives the update's existing per-step --timeout budget (1800 seconds by default), including migration and listener initialization, followed by the 12-check settle window. On the first update from an older release, the old updater invokes the newly installed CLI but does not pass that readiness budget. The candidate recognizes the existing update marker and, once the managed process is running, uses the five-minute startup watchdog instead of the standalone 60-second deadline. Migration, listener, and health transitions do not reset this bound. The old updater's subprocess timeout also remains in force. An exhausted wait reports the last observed startup phase. Standalone restarts use the progress-gated readiness wait. Verification facts and measured downtime are retained in the update run report.

When a package fails verification, the updater compares the shared and affected per-agent SQLite user_version values and configuration content with their pre-activation values. If they are unchanged and the previous runtime was verified before activation, it restores the previous package, command shim, service definition, and config writer stamp, then starts that runtime and repeats the CLI verification checks. Successful recovery leaves that Gateway running and finishes rolled-back, with the failing check kept as the reason and downtime covering service stop through verified recovery. The writer-stamp guard does not block this intentional recovery. Its allowance is scoped to rollback service commands and never persisted. See Automatic rollback. If configuration content or a schema version changed, automatic rollback is refused (state-migrated-no-rollback) and the updater attempts bounded repair on the installed candidate. The same repair slot can run when rollback itself fails. Code rollback cannot reverse state migrations. An unavailable schema comparison also prevents automatic rollback (rollback-state-unverified). After migration, a fresh candidate process finishes verification and the same durable run report. The old updater does not reopen the newer database. Git activation failures before live migrations can restore the previous source and retained built runtime. Later Git failures retain the candidate for diagnosis.

An update failure does not by itself authorize a candidate restart. Candidate activation still requires successful validation. A blocking live Doctor result does not become a restart grant. The previous runtime was verified before the update, so rollback across unchanged configuration and schemas may restart it under that prior verification and must verify it again afterward. A detached helper or Windows task autostart cannot bypass this decision.

During post-activation repair, the orchestrator starts or restarts a stopped or unhealthy service once after each turn, then reruns verification. Passing checks allow the run to succeed. If rollback already restored the previous release, successful repair finishes rolled-back and the command still exits nonzero. Otherwise it fails with the original reason and repair summaries. The agent cannot issue service lifecycle commands.

On Windows, captured Scheduled Task autostart stays suspended through Doctor finalization. The updater enables the task for activation and restores suspension if final verification fails, including after a migrated-state handoff. Native task-control failures appear in the update report. Failed suspension never triggers automatic re-enablement of the rejected installation.

On macOS, a terminated update helper can leave the selected Gateway LaunchAgent installed but unloaded and disabled across logins. openclaw doctor and openclaw doctor --fix diagnose this state. After successful standalone repair, --fix starts and verifies an already-stopped managed Gateway whose service targets the current installation. Update-time Doctor leaves activation with the updater. If the update was interrupted or installation safety is uncertain, rerun openclaw update or use Doctor and triage before starting it manually. Once verified, run openclaw gateway start (or openclaw --profile <profile> gateway start) to re-enable and start that service. Keep the same state/config and custom-label overrides. Doctor prints the selected label and recovery command. Interactive Doctor can offer bootstrap repair.

A cancellation before package mutation can restore the original service under its existing handoff ownership. Recovery succeeds only after the Gateway passes the normal restart health checks and reports the verified installation version and, for Git recovery, the exact restored build ID. A matching package version alone cannot distinguish two Git builds. A service manager accepting a start request, or reporting a live PID, is insufficient. Once the detached helper launches the updater, a missing, malformed, oversized, or interrupted direct result leaves activation to the operator. This is stricter than older helpers that restarted after an unclassified failure. Installing a new target does not change an already-running historical helper. These checks apply to the helper version that started the update.

A skipped update before activation does not park or restart the Gateway. If an interruption occurs after parking, the helper uses the child's verified recovery decision and preserves the original reason. A zero exit is retained only if required recovery succeeds or the child already verified it.

Updater exit code 79 keeps the Gateway parked only when the previous generation cannot be safely restored and verified. When the updater has restored the previous generation across unchanged configuration and schemas and supplies a verified recovery decision, the helper starts and verifies it instead of leaving it stopped. Helper recovery verifies service liveness, version/build identity, plugin activation, and channel health. It does not repeat the separate /readyz check. That report field remains unverified. The run then finishes rolled-back with the previous version and measured downtime. Missing recovery proof, migrated state, or failed restoration still requires repair before restart. A service that is observed stopped is recorded as stopped. The report does not reuse its pre-activation running status.

A terminal failed update still exits nonzero when later service recovery or triage repair succeeds. Error and skip notifications are attempted before recovery. The helper does not recreate them after the recovering Gateway consumes them. Check the final CLI result and the handoff log for the recovery outcome.

Repair the failed Doctor or installation check before restarting. Triage can inspect openclaw gateway status --deep and the update diagnostics. Avoid blindly installing older code after a newer release has migrated configuration or databases. See Updating and recovery. Restart sentinels report the outcome. Copying one does not grant permission to restart a service.

How interrupted work is detected

Persisted session status records only a run outcome: done, failed, killed, timeout, or interrupted. A new run clears the previous outcome; session rows derive running and queued from the live run registry and queue owner. Restart-aborted runs record interrupted, an interruption reason, and their end time while retaining the recovery owner's claims.

New input follows the live reply owner's queue policy. Without a live reply, durable execution fences must have terminal evidence before admission can retire them. Successful final delivery clears its pending intent and any unclaimed recovery route while preserving newer execution or delivery claims.

Before readiness, startup invokes the shared Doctor transform to normalize older persisted running or queued values to interrupted once, except verified legacy yields retain an unset outcome so their child continuation keeps ownership. It logs the repaired count, preserves recovery claims and transcripts, and retains activity timestamps. Eligible legacy running entries without a claim receive recovery custody before their old status is removed, preserving their existing automatic resume path. The saved end time is preserved or falls back to the entry's last durable update.

Archiving drains and cancels the session's work before committing the archive. Existing run outcomes, recovery receipts, and pending delivery metadata remain intact; restoring an archived session does not resume its cancelled execution. Unarchived main sessions retain their separate resume path below.

Three complementary mechanisms mark sessions whose turn did not finish:

  • At turn admission: for an ordinary text turn on an existing main session, the gateway appends the user message, clears the previous outcome, and records its recovery delivery claim in one SQLite transaction before model or before_agent_reply hook execution. Control UI does this before returning the started acknowledgement. Channel dispatch does it when the prepared turn adopts the agent run. Turns without a channel delivery route also acquire a durable recovery claim in that admission write; the claim does not grant channel delivery authority. Commands, attachments, per-turn overrides, pending deliveries, prior abort hints, plugin-owned sessions, and turns with execution hooks keep their specialized admission paths. If a before_agent_reply hook is installed, admission records enough phase state to distinguish a completed silent result from an ambiguous side-effect window. Recovery dispatches an ordinary user-triggered agent turn, so the currently loaded before_agent_reply hooks run under their normal trigger rules. Ambiguous prior hook outcomes resume with restart-safe tools rather than replaying unrestricted side effects.
  • At shutdown: during the restart drain, every session with an active run is stamped with a recovery marker in the session store before the run is aborted.
  • At startup: the gateway scans recovery-owned claims whose execution has no live owner in the new process. This catches hard crashes and kills where no shutdown code ran, without inferring recovery from a run outcome. Stale transcript lock files are cleaned up at the same time.

Agents still undergoing database startup inspection retain their recovery work. When admission finishes, the same startup recovery owner scans their stores with the original cutoff, so newly admitted turns are not mistaken for crash orphans. Inbound channel turns, including unacknowledged messages replayed after a crash, wait for their agent's writable admission before recording session metadata or dispatching. Other agents remain available, and stopping the channel or Gateway cancels the wait.

A failed store scan leaves that store eligible for the scheduled retry while other stores continue recovery. openclaw status and openclaw doctor show outstanding startup recovery failures from the running Gateway; the warning clears when the store scan succeeds.

A restart abort preserves the interrupted turn's recovery claim even when command cleanup finishes before shutdown marking. The Doctor transform also restores a missing claim for an older interrupted internal chat writer when its recorded terminal marker has no matching delivery evidence. Existing completed-turn receipts remain intact; recovery uses its normal transcript replay checks.

If a Gateway left a dead writer and an unfinished recovery cycle, sessions.recover reconciles that writer and starts a continuation in the same session. A new Control UI message also reconciles this state before admission, so a rejected send cannot trap the conversation in a "conversation changed" retry loop. Both paths preserve the session key and transcript. A live run or cloud worker still prevents this repair. Tombstoned sessions retain their separate recovery path into a new session.

If recovery fails during preparation before the agent starts, the Gateway restores the interrupted state and releases that attempt's delivery claim. The next recovery attempt uses a fresh run ID while retaining the original interrupted turn and retry budget, so a rejected pending input cannot leave the conversation permanently busy.

Automatic resume

A few seconds after startup, the gateway re-dispatches each marked session with a synthetic system message telling the agent its previous turn was interrupted by a restart and to continue from the existing transcript. If a final reply had already been produced but not delivered, its text is included so the agent can deliver it instead of redoing the work.

Startup recovery prepares and admits one continuation at a time to bound database and worker pressure. Once execution starts, the normal main lane owns concurrency; a long recovered turn does not hold a separate startup slot. Deferred database admissions join the same startup scheduler. Shutdown stops new preparation and joins the current pass, leaving unstarted interruptions available for the next boot.

Recovery follows retained transcript-window ownership when a session moves to a new key or rotates its session ID. Claim validation and cleanup read only those candidate sessions, so stale recovery claims do not load unrelated saved prompts while interactive session changes wait. Claim authority, retry budgets, stored data, and upgrade behavior are unchanged.

The restart does not cancel the user's task. The agent checks the current state, reconciles tool results whose outcomes are unknown, and continues without asking the user to repeat the request. Preparing a new message cannot consume the interruption marker; the recovery owner retains it until work is adopted or settled.

Recovery reads the interrupted turn's source before starting another run, even when a final reply is already pending. If the transcript cannot be read, the saved reply and any admitted completion claim remain available for a later attempt. Delegated requests and unverified internal inputs cannot resume automatically without surviving authority. Missing or invalid provenance does not establish a human sender for an internal claim. Legacy channel and Control UI turns retain their existing recovery checks. Child-completion follow-ups still use their existing recovery and delivery ownership checks. If their agent database is still undergoing startup inspection or preparation, a child result changes while being read, or a preparation worker refuses work at capacity, the pending completion wake retries after 30 seconds without consuming delivery attempts or changing its batch, replay identity, or retry counters. Each retry reads fresh results under the same ownership checks. The retained wake survives another restart. A confirmed inspection failure or ownership mismatch remains a failure, not permission to bypass database admission; cancellation still retires the wake.

A failed child-completion group does not stop startup recovery for unrelated groups. The original group keeps its ownership checks and retries with bounded backoff; already-started restored runs are not launched again by those retries.

When a recovered turn starts with an eligible channel delivery route, OpenClaw sends a resumption notice to that conversation, retaining its account and topic. The final reply uses the same delivery route. Transcript-only turns stay private, and a turn that has already finished does not receive a late resumption notice. A failed notice does not restart or replay the recovered work. Main-session resumption notices are best-effort and live-only: automatic-delivery permission and the recovery owner are rechecked immediately before the channel send. They are not replayed from the outbound queue; terminal-failure notice retries and normal final-reply delivery are unchanged.

Telegram renews its typing indicator while the recovered turn runs. Typing stops when the turn settles, its recovery owner changes, or the gateway closes, and respects typingMode: "never". Other channels can opt in through the guarded typing hook; unsupported channels still receive the resumption notice.

Channel plugins opt in with heartbeat.sendTypingGuarded(...). Alongside the recovery delivery target, core supplies an AbortSignal and an assertPlatformSendAuthorized callback. Plugins must honor cancellation through queued sends and invoke the callback immediately before the platform request, after asynchronous preparation. Recovery does not fall back to the unguarded heartbeat.sendTyping(...) hook.

Startup reconciliation retries transient failures up to three times with exponential backoff. Separately, each interrupted main-session cycle has a durable budget of three charged automatic dispatch attempts, retained across gateway restarts. OpenClaw charges an attempt before dispatch, refunds it when the gateway explicitly rejects the request before acceptance, and retains the charge when a post-dispatch result is uncertain to avoid replaying work. Foreground work that already owns the session keeps automatic recovery out until that work settles.

Cancelling a recovery reservation after foreground work finishes clears an empty, uncharged recovery record. Recovery reconciliation and foreground admission also clear this residue from older sessions, so it cannot block a completed subagent's requester turn. Interrupted work and pending delivery keep their recovery custody.

After the durable budget is exhausted, the session is tombstoned instead of looping forever. Inspect the failed session and use /new or /reset to start a replacement. openclaw doctor --fix can repair a stale aborted flag that conflicts with a tombstone, but it does not re-enable that recovery cycle.

If you message the failed session again in a channel, OpenClaw sends a short recovery reminder through that channel and logs each rejected message at warn level with the session key, recovery reason, and recovery command. Repeated reminders are suppressed in a bounded memory cache. Resetting or deleting the session, or restarting the Gateway, clears that suppression. Sessions with locked model selection instead direct you to Resume in new session in WebChat.

Every retry reuses one durable dispatch identifier, so an ambiguous connection failure cannot start the same recovery twice. Completed Control UI turns also retain bounded durable idempotency tombstones, allowing a reconnecting outbox to retire them without re-executing the request.

When a pending final has no remaining queue owner and its delivery outcome is uncertain, recovery records a notice for the next turn on the same route when the saved final has a delivery context and intent ID. Settling that turn clears its recovery ownership together with its delivery claim, so the next agent turn can proceed. The notice and terminal deduplication evidence survive another restart; completed work is not replayed.

Message-tool-only replies use a second durable correlation. Before a terminal same-conversation send reaches the channel, the gateway records an unresolved delivery intent on the exact session and source turn. A confirmed provider success resolves it to a durable delivered receipt. A confirmed failure clears it. Recovery completes a delivered receipt without rerunning tools. If a crash leaves the provider outcome unknown, recovery resumes with restart-safe tools so the model can inspect and report the ambiguity without replaying the external effect.

The delivered reply is also mirrored into the transcript with its source message ID. Terminal mirrors use a distinct receipt key, so a progress send with the same provider idempotency key cannot mask the terminal marker. Progress sends and receipts from older turns cannot complete the current turn. Only durable channel-ingress claims can restore message-action authority. A resumed run keeps the original source-delivery mode and source correlation, including requester identity and any same-channel/thread restriction, so the same receipt remains authoritative even if another restart happens during recovery. A message-tool-only delivery policy also survives run admission and model fallback; private final text never becomes a pending automatic reply. A message-tool-only turn without reconstructable channel authority is tombstoned because OpenClaw cannot safely mint message-action authority without the original channel-ingress claim. The terminal notice directs the user to start a replacement with /new or /reset.

A recovered Control UI turn can finish pinned dashboard widgets using the interrupted turn's exact session and recovery claim. A browser connection does not survive the restart: recovery retains dashboard authoring, while inline and device presentation still require their normal client capabilities.

Before resuming, the gateway classifies the transcript tail to choose the tool restriction for the continuation. An aborted turn is the interruption itself, so it resumes on a best-effort basis whatever abort detail the provider or worker recorded with it: partial streamed text stays in the transcript and the continuation picks up from the message beneath it, while a tool call left dangling is dropped from the next provider payload. Provider failures, completed assistant tails, empty transcripts, and stale pending approvals also continue from the existing transcript. States with ambiguous side effects normally use restart-safe tools. A session with effective Full Access, including an inherited Full Access default, keeps its ordinary tools so it can inspect the outcome and finish the task. Recovery does not replay the interrupted call automatically or treat its missing result as success. Existing tool restrictions and current permissions still apply. Recovery prompts identify interrupted, missing, or aborted tool results as unknown outcomes from the Gateway restart. A follow-up to an interrupted native child receives the same context so it can verify effects before retrying a tool call. Pending reply delivery, ambiguous reply-hook outcomes, and explicitly replay-safe Code Mode reconstruction retain their narrower recovery restrictions.

OpenClaw can also reconstruct interrupted read-only Code Mode work. Code Mode marks these runs as restart-safe and rejects side-effecting catalog or namespace tool calls before they execute. If a restart lands on the wait control, the new gateway reconstructs the turn from its transcript and forces the reconstructed execution to remain restart-safe even if the model omits or clears that flag. The host filters the entire reconstructed turn to audited read-only core tools and explicitly replay-safe plugin tools, including when Code Mode is disabled after the restart. Other interrupted Code Mode work resumes for model reconciliation: Full Access keeps its configured tool surface unless the current turn has an explicitly replay-safe checkpoint. Other sessions retain the restart-safe restriction. Old process-local runs and approval handles are not revived.

Subagents

Subagent runs are persisted in the shared SQLite state database, so the subagent registry survives the process. On boot, interrupted child runs settle through their normal completion path as soon as startup restores requester ownership, without waiting for the periodic registry sweep. The sweep remains a retry backstop. The crash-loop breaker pauses this settlement too; the same sweep retries when the breaker's recovery window ends. They are not automatically relaunched. The parent receives the interruption outcome and owns finishing the user's task. Its recovery input lists current unfinished child session and run identities, including children interrupted by the restart. Older runs superseded by a newer child run are omitted, as are records from another store, parent session, or parent lifecycle revision. A reset keeps the session ID but changes its lifecycle revision, so retained child work from before that reset cannot enter the new parent's actionable recovery roster. Large lists show the first 32 children and tell the parent to inspect the remaining children.

The parent must reconcile each unfinished child with its saved history and the original request. When the child's task is still needed, it should prefer a follow-up in that retained session with sessions_send. It first confirms that the old execution stopped and checks any uncertain tool effects. It can use work already completed, assign a replacement, or finish the remaining work itself. Recovery does not automatically replay child commands or duplicate running work. An interruption alone is not a blocker; the parent continues until the request is finished or a specific blocker requires user input or unavailable authority. Existing cleanup and retention settings still apply.

Startup retires superseded requester completion claims through the registry's cleanup owner before restoring the surviving claim. Those historical rows do not block current children in the same requester turn or recovery of other subagents. Saved yield intent is evaluated from the same current children used for the transfer. Interrupted current children still report their restart outcome to the parent. Turns with only superseded children need no requester settlement.

If a parent yielded while waiting for children, its saved batch collects both completed and interrupted results and wakes the parent once the batch settles. Yielded turns leave the persisted outcome unset while their continuation still owns unfinished work. A parent already working on those results resumes through ordinary main-session recovery. A child result or an announce: run identifier does not make unfinished parent work disposable. Both recovery paths give the parent the interrupted children's identities and the same reconciliation guidance. Restart interruption remains in history as an interrupted outcome, rather than a child execution failure. Genuine execution and delivery failures still require attention.

Saved batch wakes keep their original batch identity and completion-delivery contracts. Their actionable recovery roster includes only children whose captured parent ownership still matches. Older or stale records remain ordinary completion history; they cannot add new instructions to continue those child sessions. Current reset already revokes its live saved wakes. Recorded outcomes and historical child metadata are retained.

On upgrade, saved interruptions that retain the typed restart-recovery owner are reconciled through the same startup path. Historical failed runs without that ownership evidence remain failed: their error text alone cannot distinguish a restart from a genuine failure. Inspect those retained sessions before continuing them; recovery does not rewrite ambiguous history.

A completed child may still owe its requester a final follow-up. If that follow-up is waiting to retry or is interrupted by restart, the saved obligation survives and resumes after startup. Restart admission rejection does not consume an attempt, and cancellation of an admitted attempt does not exhaust the obligation. Existing delivery retry limits still apply. Transient cleanup preparation failures keep required final delivery scheduled within its existing delivery window. Incidental cleanup retains its bounded retry limit. Settling a yielded turn's wake leaves its unfinished native run and final delivery intact. Completed cancellation keeps its wake and cleanup bookkeeping in the native subagent record; it does not require a separate Tasks row. Recovery reconciles an expired cancellation's retained marker before retrying its requester wake, preserving the original cleanup record. Live child cancellation can wake a waiting requester while normal cleanup reconciliation completes. A delayed cancellation callback cannot reopen completed cleanup.

Agent-requested restarts

When the agent itself triggers a restart (applying a config change, updating the gateway, or an explicit restart request), a restart sentinel is written to SQLite before the process exits. After boot the gateway posts the outcome back to the originating chat and dispatches any requested one-shot continuation turn so the agent picks up exactly where it left off, on the same channel and thread.

For updates, the sentinel carries stats.runId, linking the detached updater to its durable update_runs record. The new Gateway records its observed running version, build, and startup facts there. It preserves a terminal outcome already written by the updater and waits while a managed handoff is still pending. Before preparing notices or continuations, it reconciles a newer final sentinel for that same run and handoff. A pending sentinel keeps its existing bounded retry window even when the ledger is already terminal; an unrelated replacement remains untouched. If the helper never publishes its final sentinel, expiry reports the recorded terminal outcome without changing it. A still-running Gateway-owned row finishes as failed with restart-unhealthy; CLI-owned runs retain their updater's authority and outcome. The post-restart notice is rendered from that row using the same report as openclaw update status. Consuming the sentinel does not remove run history. Sentinels left by older releases retain their existing delivery route.

Any update run with an existing internal origin session, including Control UI and webchat, appends its report directly to that session's transcript, even when the caller supplied only sessionKey and no deliveryContext. A completed update with no continuation does not wake the model to deliver the report.

The sentinel's typed SQLite columns are authoritative for restart handling. Its payload_json value is a replay/debug shadow only. Runtime reads, writes, and clears SQLite state without a file fallback. The pre-July restart-sentinel.json format and its .doctor-importing claim are retired. Doctor, update admission, and restart recovery leave these files and any current SQLite notification unchanged, and require an intermediate upgrade through 2026.9.7 with openclaw doctor --fix before retrying. Supported updaters, including 2026.6.34 and 2026.9.2, write native SQLite state and retain their existing restart delivery and continuation behavior.

Safety valves and observability

  • Crash-loop breaker: 3 unclean boots within 5 minutes trip a breaker that suppresses auto-start side services on the next boot, so a crashing gateway does not amplify itself. A continuously stable safe-mode gateway rechecks the breaker after the full unclean-boot window drains and then resumes deferred channel auto-start without requiring another gateway restart.

    When the breaker is tripped, the control plane still starts, but channel plugins (and other auto-started side services) stay down until an operator manually overrides the suppression or the full window drains with no unclean boots. Recovery preserves channels that an operator manually stopped and any separate development-mode suppression. Gateway logs look like: channel autostart suppressed by crash-loop breaker; refusing automatic start for <channel>… Start a channel manually with: openclaw gateway call channels.start --params '{"channel":"<id>"}'

    Safe mode also pauses main-session restart recovery. Interrupted entries and transcripts stay unchanged: the pause does not charge an attempt, settle a turn, send a notice, or tombstone a session. The recovery log names the pause and the breaker window end. Once the full window drains, the same healthy process automatically re-arms its existing recovery scheduler and resumes the interrupted work once, using the original startup cutoff and ownership checks. A new unclean boot extends the pause; shutdown or lifecycle replacement cancels the pending re-arm. No manual retry or extra Gateway restart is required.

    Operator recovery SOP:

    1. Confirm the gateway process is up (openclaw gateway status / LaunchAgent or systemd unit still running). A “channel disconnected” symptom often means suppressed autostart, not a dead gateway.

    2. Inspect channel state: openclaw channels status (add --probe when useful). Look for stopped / not connected accounts while the gateway itself is healthy.

    3. Fix the root cause of the unclean boots (bad config, plugin crash on start, missing secrets) before forcing channels back up.

    4. Manually start a channel while suppression is active:

      bash
      openclaw gateway call channels.start --params '{"channel":"<id>"}'# optional: {"channel":"<id>","accountId":"<account>"}

      channels.start is a manual override. It does not disable the breaker for other channels.

    5. Or leave the healthy gateway running until the full unclean-boot window drains. The same process logs that the restart-loop breaker recovered and starts the deferred configured channels. If that message does not appear after the window plus one health-monitor interval, inspect the gateway logs and run openclaw doctor before restarting.

    See also Gateway (safe mode paragraph) for the same control-plane vs channel-autostart split.

  • Main-session attempt budget: three charged automatic dispatch attempts per interrupted cycle. Exhaustion tombstones that session until it is inspected and replaced.

  • Metrics: recovery activity is exported via Prometheus as openclaw_session_recovery_total and openclaw_session_recovery_age_seconds.

  • Logs: recovery decisions are logged under the main-session-restart-recovery and agents/subagent-registry subsystems. Every startup pass includes bounded skip counts by reason such as live_owner, work_start_blocked, or dispatch_target_unavailable, even when other sessions started. Each interrupted main candidate has a structured decision line with boot/pass, session and source-run identity, outcome (started, settled, deferred, or blocked), and the next responsible owner.

  • Reply hooks: resumed turns run currently loaded before_agent_reply hooks under the normal user-trigger rules. Automatically delivered replies also run the normal reply_payload_sending hook before channel delivery, with the recovered session, run, account, and conversation context.

Verify recovery after an update

A healthy Gateway confirms availability, not completion of interrupted work. Check each previously active session and its child tasks: it should have finished before shutdown, resumed execution, or reached a visible terminal outcome or recovery error. Check queued inputs separately for transcript consumption or an explicit unresolved state.

The main-session recovery log distinguishes execution resumption from the later main-session restart recovery terminal event. A recovered count at startup means execution resumed. It does not prove that an assistant reply was delivered. Use the session transcript and recorded delivery outcome to verify completion.

What is not resumed

  • Sessions excluded from main-session recovery because another owner already handles them: subagent sessions (settled back to their parent), cron sessions (the scheduler re-runs on schedule), and ACP-managed sessions (the connected IDE or client owns the resume).
  • Work that was never admitted: messages arriving during the drain window are rejected with an explicit restart error rather than silently queued into a dying process.
  • Gateway terminal PTYs, including operator- and agent-owned terminals. They are process-local and end when the Gateway restarts.
  • Standalone embedded turns cannot take over a main session with pending restart recovery because they do not share the gateway's lifecycle owner. Run the turn through the gateway or reset it there with /new or /reset.
Was this useful?