Agent Auto-Restart¶
If a sase update breaks a running agent before it did any model work, sase puts it back once, under the same name, and tells you exactly what it did and why.
An automatic restart is ,x plus an unmodified submit, run headlessly. Each lineage
gets one automatic restart. Every other update-shaped failure is surfaced with its
reason and never silently swallowed.
When a restart happens¶
Three things must all hold before sase relaunches anything:
- Signature. The error belongs to a known update-skew family: a torn
ImportError/ModuleNotFoundError/ moduleAttributeErrorin sase's own code, a stalesase_core_rsbinding or wire schema, or a data-format skew ("written by a newer sase version"). A matching error pattern alone is never enough.TypeError/NameErrorsignature mismatches are always treated as real bugs and only annotated. - Witness. At least one independent witness must corroborate the update: the
runner's boot code identity no longer matches the current tree (W1), a
sase updatejournal row falls between boot and the failure (W2), or file-level proof names the culprit commit (W3). - Probe. A fresh interpreter must import the current versions of every managed module on the failing traceback (W4), proving the new tree is healthy and the death was the swap, not the code.
The healer also waits for quiescence: no update holds the code-swap writer lock and
the tree has been quiet for quiescence_seconds. A pass that cannot act yet defers (up
to max_defer_seconds) instead of giving up.
Phases and modes¶
Only deaths before the model turn are relaunched (booting, waiting,
preparing):
| Death phase | Mode | What happens |
|---|---|---|
| Before the model turn | relaunch |
Relaunched once under the same name |
After the model turn (provider_running and later) |
notify_post_provider |
Not relaunched; you are notified with held-workspace guidance |
| Plan, question, monitor, gate, or pipe handoff | ask |
Not relaunched; you are asked to decide |
Never restarted: provider errors, rate limits, auth, kills and cancels,
OOM/timeout/disk-full, directive or macro errors, tool and test failures, and any
ImportError from workspace or third-party code. Anything the restart planner refuses
is also left alone.
Skipped on purpose, every time: killed or dismissed agents, agents mid-,x, and agents
holding a question or gate. User intent wins.
The ledger and the one-restart guarantee¶
Claims live outside every artifacts directory under ~/.sase/agent_auto_restart/ so a
workspace wipe can never remove them:
ledger/<project>__<lineage_root>.json— one record per lineage, claimed atomically before any mutation.doorbell/<project>__<artifacts_timestamp>.json— dropped by the dying runner; deleted once a ledger record owns the failure.episodes/<episode_slug>.report.json— the live per-episode report, re-rendered on every change.state.json— storm-breaker pause state.
The state machine is claimed → deferred | declined | launching, then
launching → launched | settled_failed and launched → settled_ok | settled_failed. A
replacement that fails again is reported and never retried. When a claimer died
mid-launch, a later pass adopts the replacement it finds or settles the record as failed
— it never launches twice, because forced name reuse would wipe the replacement it just
started.
Update episodes are grouped by culprit commit (for example sase@9fd8a08), so one
update that breaks five agents produces one story, not five.
CLI¶
A bare sase agent auto-restart lists ledger records, newest first, grouped by episode.
| Command | Purpose |
|---|---|
list [-a/--all] [-j/--json] |
Ledger records, newest first, grouped by episode |
resume |
Re-arm after the storm breaker trips |
run (NAME \| -a/--artifacts-dir DIR \| -p/--pending) [-n/--dry-run] [-j/--json] |
Run the healer; -p is the scheduler job's target and a manual run still honors the ledger |
scan [-j/--json] [-l/--limit N] [-s/--since DURATION] |
Read-only replay of the classifier over history; never writes the ledger |
show TARGET [-j/--json] |
One card: verdict, witness checklist, state timeline, evidence paths |
Configuration¶
agent_auto_restart:
enabled: true # permanent kill switch
quiescence_seconds: 30 # tree-quiet wait before acting
max_defer_seconds: 1800 # longest a claim waits for quiescence
pending_resurface_seconds: 600 # stale pending rows are re-surfaced loudly
storm_max_per_episode: 12 # storm breaker: launches per episode
storm_max_per_30m: 20 # storm breaker: launches per 30 minutes
Set enabled: false to turn the feature off permanently. While off or paused, pending
failures are re-surfaced loudly instead of healed, and the ledger is never written —
disabling never swallows a failure and never spends a lineage's one restart.
Notifications¶
One upserted amber ↻ row per update episode (sender=agent.auto-restart,
action=ViewReport). The first relaunch creates the row and its single information
toast; every further relaunch in the episode appends one +1 note without toasting and
refreshes the title and the inline report snapshot in place. The title never contains
"fail" or "error"; exception text appears only in note 2 or later.
The live report (headline, update/culprit/witness facts, per-agent rows with a live
Now column, left-alone bullets, why-text, prevention line) is linked from the row.
Escalations are loud and land in the Errors bucket: post-provider deaths name the held
workspace, declined relaunches name the reason, re-broken replacements say sase will not
retry, and a tripped storm breaker pauses the feature until
sase agent auto-restart resume.
In the Agents tab, in-flight recoveries render as amber ↻ RESTARTING with a dim reason
instead of red FAILED. The replacement keeps the same name with a ↻ chip and a
provenance block naming the update, the signature, and the preserved error_report.md
(opened with the existing v key).
Troubleshooting¶
auto-restart never ran — is the sase scheduler running?The row satpendingpastpending_resurface_secondswith no ledger record. Start the scheduler; nothing was claimed or spent.- "Couldn't restart \
automatically — \<reason>." The classifier, probe, or quiescence gate declined. Press,xon the row to retry by hand. - "This was its automatic restart — not retrying." The replacement broke again. Its one restart is spent; investigate it as a real failure.
- "Auto-restart paused: N agents broke within one update." The storm breaker
tripped. If the update really is that breaking, run
sase agent auto-restart resumeto re-arm. sase agent auto-restart scan -s 120dreplays the classifier over history read-only. Zerorelaunchverdicts among non-skew failures means the classifier is not over-eager.- Evidence. Every relaunch preserves the error report, traceback, log tail, facts,
and verdict in the recovery bundle linked from the ledger record (
show TARGET) and the episode notification.