Skip to content

Development

This page orients contributors working in the sase repository. It covers local setup, verification, source layout, and documentation publishing paths.

Setup

Requirements:

just install-venv
sase --help

just install-venv installs the package in editable mode with development dependencies. When a sibling ../sase-core checkout is present and cargo is available, it also builds and installs the local sase_core_rs extension before resolving Python dependencies.

The verification recipes cache their setup-validation verdicts inside the active virtual environment. The cache is fingerprinted from pyproject.toml, uv.lock, the validator implementations, the local sase-core version, and the installed environment metadata, so dependency or environment changes revalidate automatically. Set SASE_TEST_SETUP_FORCE_REVALIDATE=1 on any just invocation to bypass the cache while diagnosing setup problems.

Your sase versus this checkout's .venv

Two different environments matter when you hack on SASE, and mixing them up is the most common setup mistake:

  • Your sase command lives in the uv tool environment. It is what runs when you type sase anywhere, and it is what every agent and the scheduler import.
  • This checkout's .venv (created by just install-venv) is where tests, lint, and benchmarks run. It never affects your sase command.
just install       your `sase` command ← the latest PyPI release
just install-dev   your `sase` command ← this checkout + its paired sase-core
just install-venv  this checkout's .venv ← tests, lint, benchmarks (never your `sase`)

As a contributor you almost always want just install-venv: edit the checkout, run just check, and the .venv picks up your change. Reach for just install-dev only when you want your global sase command itself to run this checkout (dogfooding unreleased code). It installs the checkout editable plus its paired sase-core — $SASE_CORE_DIR, else <sase checkout>/../sase-core, the same checkout sase update rebuilds from — so it needs git and cargo on top of uv. Day-to-day updates after that belong to sase update, which maintains the install in exactly the shape install-dev creates. See Installing from a checkout for the full comparison.

just install and just install-dev are human-only: inside a SASE agent or monitor they refuse with exit 2, because they would replace the sase binary every running agent depends on. The refusal points at just install-venv for workspace repairs. If the human genuinely asked for a global reinstall, the agent proposes it through /sase_gate, and the approved run carries a reason in SASE_GLOBAL_INSTALL_BYPASS='<reason>' (the installer prints a one-line note naming that reason). Agents never set the bypass on their own authority.

Verification Commands

just install-venv  # Install with dev deps
just fmt           # Auto-format code and Markdown
just fix           # fmt plus keep-sorted fixes
just lint          # Run ruff, mypy, repository audits, symvision, toobig, and keep-sorted
just test          # Fast parallel test run, excluding slow and PNG visual snapshot tests
just test-cost     # Fast suite with cost attribution and committed budget checks
just test-slow     # Slow pytest subset only
just test-visual   # check-only alias of `just fix-tui-screenshots --check`
just fix-tui-screenshots  # capture, compare, and apply ACE and pager PNG goldens
just update-visual-snapshots  # update alias of `just fix-tui-screenshots`
just test-terminal-smoke  # Optional real-terminal sase's TUI smoke test
just test-cov      # Parallel test run with coverage + 50% gate, excluding visual snapshots
just test-contexts # Record the per-test coverage baseline the selector consumes, and cache it host-locally
just test-ace-page-group-isolated  # Rerun AcePageGroup modules with fresh AcePage checkouts
just test-contention  # Diagnostic soak: repeat the default lane under pinned-CPU contention and tally per-node failures
just check         # Agent default: whole-repo lint and validation gates + a diff-scoped test lane
just check-full    # Exhaustive verification: whole-repo gates + the full test suite + flake baseline + local screenshot update
just validate      # sase-core-rs minimum, static feature flags, and sase validate
just selection-health  # Health of the diff-scoped test lane, including false negatives
just selection-backtest  # Replay real history and measure selection recall against coverage
just refresh-contexts-baseline  # Cache CI Telemetry's per-test coverage baseline for selection
just refresh-contract-manifest  # Regenerate tests/contract_manifest.txt from the marker
just refresh-shard-timings  # Refresh tests/shard_timings.json, the master gate's shard balance table
just test-tox      # Test across Python 3.12, 3.13, 3.14
just clean         # Remove build artifacts
just build         # Build wheel and sdist

Diff-scoped checks (just check)

just check is the agent default: every whole-repo lint gate except toobig runs, followed by just validate (the published sase-core-rs minimum, static feature-flag checks, and sase validate), committed plan validation, and an advisory probe of the published sase-core-rs floor, but the test stage is just test-scoped instead of just test. Agents can rarely act on toobig failures, and the toobig_split routine owns those splits; the gate still runs in just lint, and therefore in CI. just check-full runs the same gates. tools/select_tests builds a cached import graph from src/** and tests/**, seeds it with the changed and untracked files in the current diff against $SASE_CHECK_BASE (default origin/master), and walks reverse import edges out to a bounded depth (SASE_TEST_SELECTION_DEPTH, default 2) to find the test files that plausibly exercise the change. The selection always includes the curated contract set (tests/contract_manifest.txt) and excludes tests/ace/tui/visual/** unconditionally. The scoped run is serial (-n 1) unless the middle gear below wins it a small lease, and it never queues behind other agents' runs either way.

Selection is a heuristic, not a guarantee: an unbounded closure would select the vast majority of the suite because of a large import cycle in src/sase, so depth-bounding is the mechanism, not a tuning knob. A handful of broadening rules escalate to the full suite when the change touches something the closure cannot safely reason about (a conftest, pyproject.toml, the Justfile, config schemas, the selection engine itself, or a narrow set of environment-identity inputs — see "The core-identity-changed escalation" below). A selection that survives those rules is then costed rather than counted: it escalates when a serial run of it is estimated to take longer than SASE_TEST_SELECTION_MAX_SERIAL_SECONDS (default: the full lane's measured wall clock, 444s), and only where no such estimate is available does the file-count ratio SASE_TEST_SELECTION_MAX_RATIO (default 0.25) decide instead. An escalated run falls through to the same governed, fully parallel lane as just test.

The runtime budget exists because file count is a 6x-spread proxy for runtime: measured on athena on 2026-08-06, eight of 39 scoped runs took longer than the then-232s full lane and consumed 75% of the lane's total wall clock, and the worst — 494 files, which the ratio rated as scoped — ran 1,032.6s where just check-full would have finished in ~291s. Recalibration on 2026-09-14 found the current governed full lane at 14 workers runs in 435-493s, with a 443.5s median across five recent full fast records, so the default crossover is 444s. Past the crossover the fast path is the slow path, so the lane stops taking it. Every scoped manifest records both halves of the comparison (max_serial_seconds and the timings block), and tools/select_tests --explain prints them whether or not the rule fired — serial budget: estimated 180s against a 444s budget (within; 96% of the selection covered by the timing table).

Agents run just check, not just check-full. just check-full is the exhaustive local lane — every lint gate except toobig, the full suite through just test-cost, the flake-baseline gate, and a local TUI screenshot update — and agents invoke it only when the current prompt, the user, or the assigned bead explicitly names that command (typically a CI failure on a check-full-only gate). Landing, touching the broadening set, and a scoped escalation are not reasons to start it; just check already escalates internally when the selector cannot trust the closure. CI always runs the full non-visual suite and a dedicated check-only visual job, so a scoped false negative surfaces there within roughly the CI test leg's runtime; it is the backstop, not a silent gap. A just check pass with a just check-full failure is a test-infrastructure bug; file it rather than treating it as remaining product work.

Both recipes are guarded: in a SASE agent process (SASE_AGENT set), just check and just check-full refuse with exit 2 unless they run inside sase tool run check / sase tool run check-full for this checkout, or with an explicit SASE_TOOL_BYPASS='<reason>' (the guard also fails open when sase is not on PATH). Humans, CI, finalizers, monitors, and procs run them raw as before.

sase tool run check        # how an agent runs `just check`
sase tool run check-full   # only when explicitly instructed

Use tools/select_tests --explain to see which rules fired and why a given file was pulled into (or excluded from) the current selection. The selection manifest — the resolved base, changed files, rules fired, and selected test files — is written to .pytest_cache/sase-selection/manifest.json on every scoped run.

Per-test-file timings

File count is a poor proxy for how long a selection takes to run: measured on athena on 2026-08-06, a 94-file selection ran 465s serially while a 517-file one ran 404s. So the lane also measures cost. Every full-lane run and every scoped run loads tests/_test_selection_timings_plugin, which sums each test's setup/call/teardown wall seconds up to the test file and writes them to ${SASE_HOME:-~/.sase}/test-selection/<project-key>/timings/. One just test covers every test file in a single pass, so the table bootstraps from a run that already happens; a scoped run then refreshes the files the lane touches most. The newest eight recordings are merged newest-wins and the rest pruned.

tests._test_selection_timings.estimate_serial_seconds() turns that table into what a serial run of a given selection would cost. It never guesses silently: files the table has not seen are extrapolated at the covered files' mean only while at least SASE_TEST_SELECTION_TIMINGS_MIN_COVERAGE (default 0.8) of the selection is covered, and below that the answer is an explicit "insufficient data" rather than a number. Every scoped manifest records the estimate, the coverage fraction, and the identity of the table it came from, under timings.

The estimate is what the serial-budget-exceeded rule above decides on. Where it is unavailable — a fresh host, a mostly-new selection, or SASE_TEST_SELECTION_TIMINGS_DISABLED=1 — nothing changes: the file-count ratio decides, as it did before the table existed. SASE_TEST_SELECTION_TIMINGS_DIR relocates the table.

Suite-cost budgets

just test-cost runs the same marker selection as just test, but it loads tests/_test_cost_plugin.py, writes a cost recording under the timing store's cost/ subdirectory, prints tools/test_cost_report, and then enforces tests/perf/baselines/test_cost_budgets.json with tools/check_test_cost_budgets. just check-full uses this lane so the exhaustive local path catches cost regressions; ordinary just test, just test-cov, and just check keep the lower-overhead timing recorder. CI enforces the same budget on the Python 3.13 full-suite leg. Because just check-full runs this lane instead of just test, the cost lane also records full-run failures for selection health; it deliberately does not record per-test-file durations, because the probe taxes exactly the numbers that table is for.

The report has three parts: summary totals, cause attribution, and top files. The summary budgets guard total per-test wall seconds, idle seconds, collection seconds, and peak worker RSS. Cause budgets guard the hot buckets this suite has historically regressed: sase's TUI app/page startup, full parser builds, YAML reparses, and avoidable subprocess round-trips. Local runs use a narrower tolerance than CI so host noise does not make shared runners brittle.

When a budget fails, first run just test-cost -- <path-or-node> around the suspected area and compare it with the committed baseline:

just test-cost -- tests/main/test_parser.py
tools/test_cost_report --top 20

New tests should treat bare pilot.pause(), positive fixed sleeps without an inline # sase-test-wait: <reason> pragma, one-app-per-assertion sase's TUI boots, full create_parser() builds when a narrower command tree is enough, and CLI subprocesses used only to inspect stdout as defects. Use observable wait helpers, shared or cached test helpers with explicit reset semantics, and in-process entry points unless the process boundary is the behavior under test.

The middle gear

The lane used to have two gears: one worker with no lease, or the whole governed suite. serial-budget-exceeded is the one escalation a third gear can answer — the selection is sound, it is merely too slow to run serially — so when that rule fires alone, tools/run_pytest asks the suite gate for up to SASE_TEST_SELECTION_SCOPED_WORKER_CEILING (default 4) worker tokens and runs the selection at whatever width it gets. The 494-file selection that ran 1,032.6s serially was ~258s at four workers, against the historical 232s full-lane crossover that first motivated the gear.

The request is a single non-blocking attempt (WorkerTokenLease.try_acquire). If the tokens are not free right now, the run escalates exactly as it did before — the lane never queues behind another agent's run, which is the property the whole scoped lane exists to deliver. The gear also declines a one-token grant (a serial run wearing xdist's bookkeeping), a run that must stay serial anyway (--inline-snapshot=fix/review), and any run already under a governed parent (SASE_TEST_GATE_DISABLED=1). Setting the ceiling below 2 turns the gear off entirely.

The width is the gate's decision, not the caller's, so scoped mode still rejects -n and SASE_PYTEST_WORKERS. A change-set escalation — a conftest, the Justfile, core-identity-changed — is never offered to the gear: those rules fire because the closure cannot be trusted for that change, and no amount of parallelism answers that.

Every run that reached the gear records a gear block on its manifest (granted, width, ceiling, and the refusal reason when there is one), and it shows up in just check's scoped summary line as gear 4 workers or gear refused (tokens-unavailable). A granted run is recorded as not escalated, because it ran a selection: its real duration and its width both reach the health store, where just selection-health counts the gear's runs and refusals and reports the width mix behind the duration percentiles.

The core-identity-changed escalation

tools/validate_test_environment already digests the installed environment to invalidate its own validator-verdict cache — pyproject.toml, uv.lock, the venv's pyvenv.cfg, the sibling sase-core/Cargo.toml, its own four validator scripts, every installed distribution's dist-info metadata, the compiled sase_core_rs extension, the venv's bin/python, and a core-source bucket (the linked sase-core checkout's source identity plus the venv's built-from stamp — see Source / development workflow). The selector reuses those same digests rather than forking a second, divergent fingerprint, but as a per-input map (tools/validate_test_environment._fingerprint_inputs) instead of one opaque combined hash, so it can say which input moved instead of only that the environment did.

Only some of those inputs are worth forcing the whole suite over. core-identity-changed fires only when a bucket in tests._test_selection_manifest.ENVIRONMENT_ESCALATING_INPUTS differs from the previous scoped run's manifest: pyproject, uv-lock, venv-config, core-cargo, extension, and python. Each is either invisible to git diff against the merge base (the sibling repo's Cargo.toml, the compiled extension, the interpreter identity) or covers a case the diff-visible packaging-config rule cannot: pyproject.toml/uv.lock already broaden the selection via packaging-config when they are part of the current diff, but this fingerprint also catches the same files changing between runs without being part of it — a git pull that lands a dependency bump the working diff never touches. The four validator:* scripts and every installed package's metadata (environment-metadata) are recorded for attribution but do not escalate on their own: they are repository tooling and environment bookkeeping, not something that changes which tests exercise the diff. core-source does not escalate either: when the linked source moves, _setup rebuilds the extension, and the rebuilt extension's extension bucket escalates instead. A run where only a non-escalating bucket changed falls back to the normal closure plus contract-set-always, the same as any other unremarkable diff — not to silence, since the manifest's baseline.environment_changed_inputs still lists every bucket that moved, escalating or not, and tools/select_tests --explain prints it as environment inputs changed: ... whenever it is non-empty.

The compiled extension's identity was previously untracked in practice: sase_core_rs installs to the nested site-packages/sase_core_rs/sase_core_rs.abi3.so, but the old glob (sase_core_rs*.so applied directly to site-packages) does not cross the /, so it matched nothing and the extension input was silently empty — a rebuild was caught only indirectly, through the dist-info METADATA version. It now searches site-packages and site-packages/sase_core_rs, mirroring tools/purge_sase_core_rs_extensions's candidate directories, and hashes the file's content instead of its stat(), so a rebuild that reproduces identical bytes is not a change.

Measured on the epic's research (2026-08-06, 63 scoped runs against the real host store): core-identity-changed fired in 16 of them, and was the sole reason for escalation in 8. The single-digest scheme those runs recorded could not say which input caused any of them — that attribution is unrecoverable for those historical runs, which is exactly the gap the per-input map above closes for every run recorded from here on.

Coverage-context ground truth

The import graph cannot see dynamic dispatch, plugin lookup, or config discovery. Per-test coverage can. CI Telemetry's coverage-contexts job runs the fast suite with --cov-context=test, so its .coverage database records which test executed each line, and publishes it as the sase-coverage-contexts-<sha> artifact.

That job is deliberately separate from the per-PR coverage leg, and runs in scheduled CI Telemetry master jobs only. Measured on athena on 2026-08-06 over the full fast suite at 12 workers:

Variant Suite runtime .coverage gzipped
branch coverage (the PR leg today) 470s 17 MB 7.7 MB
branch coverage + contexts 538s 906 MB 283 MB
contexts, line coverage only 474s 49 MB 12.1 MB

Branch coverage stores every arc per context; line coverage stores one bitmap per (file, context). Selection only ever asks "which tests executed this line", so coverage_contexts.toml turns branch coverage off and the PR leg keeps its branch data and its 50% gate untouched. Baselines are resolved as ancestors of an agent's HEAD, so a per-PR database would be one nobody ever looks up.

just test-contexts                      # record a baseline locally (what CI Telemetry runs) and cache it
just refresh-contexts-baseline          # newest master baseline that is an ancestor of HEAD
just refresh-contexts-baseline --force  # re-download even if already cached

There are two supply routes, and neither is a network dependency at selection time. The artifact is published by scheduled CI Telemetry runs and retained 14 days, so a host that has been idle longer than that — or is offline, or never fetched — would otherwise run the scoped lane on the static closure alone. just test-contexts closes that hole: on success it runs tools/install_coverage_contexts, which files its own .coverage in the cache as <HEAD sha>.sqlite. Because the cache is host-local rather than per-workspace, one instrumented run in one numbered workspace supplies every workspace on the machine. Instrumentation stays opt-in — nothing on the just check or just check-full path records contexts — and SASE_TEST_SELECTION_INSTALL_CONTEXTS=0 records without caching.

cov-contexts runs pin COVERAGE_CORE=ctrace. On Python 3.14 coverage otherwise defaults to the sysmon core, which stops monitoring a code location once it has been seen — so only the first test to execute a line is credited with it, and per-test attribution thins out as the suite runs. Measured on athena at 6b0976bcb: over the full suite, tests/test_agent_lanes.py recorded 6 contexts against agent_lanes.py under sysmon and 32 under ctrace, which is what CI's Python 3.12 leg (already on ctrace) records. A local baseline has to be the same ground truth, not a thinner one.

The installer refuses three databases that would be worse than no baseline at all, since a baseline that resolves but contributes little silences context-baseline-missing while adding few tests: one recorded against a src/ tree with uncommitted changes (its line numbers are not the commit's line numbers), one recorded over part of the suite, and one whose attribution density — (file, test) pairs per measured file — is under half that of the densest database already cached. The third guard is the one the other two cannot see: a sysmon-cored run names the whole suite over a clean tree and still holds an order of magnitude less ground truth. --allow-dirty, --allow-partial, and --allow-thin override them deliberately; a refusal never fails the recording recipe.

Baselines are cached by SHA under ${SASE_HOME:-~/.sase}/test-selection/contexts/, newest five retained however they arrived, each beside a <sha>.sqlite.breadth.json sidecar recording the context, attribution, and file counts its producer measured. Selection itself never touches the network: among the cached ancestors of HEAD it reads the nearest one that is not materially thinner than the broadest available, and an absent or unreadable one is not an error — the run records context-baseline-missing and proceeds on the static closure alone, so a fresh workspace with no connectivity still gets a working just check.

Breadth is what breaks the tie, not recency. Ranking on file mtime held while every baseline arrived the same way, as a CI artifact, and stopped holding once a local run became a second producer: measured on athena at b08862001, a local 6b0976bcb database (14,349 contexts, 46,364 attribution pairs) outranked CI's 96183d71b (58,770 and 597,959) purely by being written more recently, over a near-identical file count. Every selection that resolved it got 13× less attribution while reporting a healthy context-selection. So the cache now ranks ancestors by breadth first and commit distance second, with anything holding at least 75% of the best candidate's attribution pairs counted as comparable — a gate wide enough that ordinary run-to-run variation still lets the nearer baseline win.

Contexts are unioned into the selection, never substituted for it. They are ground truth only for the code that existed when the baseline was recorded; they say nothing about code added since, and a brand-new test file has no context rows at all. A baseline more than SASE_TEST_SELECTION_CONTEXTS_MAX_DISTANCE commits behind HEAD (default 50), or one whose commit this workspace does not know, is still used but records context-baseline-stale so just selection-health's rule histogram can show whether staleness correlates with false negatives. Set SASE_TEST_SELECTION_CONTEXTS_DISABLED=1 to ignore the cache entirely, and SASE_TEST_SELECTION_CONTEXTS_DIR to point it elsewhere. The manifest's contexts block records the baseline SHA, its distance behind HEAD, whether it was stale, which changed files it matched, and how many test files it contributed.

Contexts are consulted only on the path that actually produces a narrowed selection. A run a broadening rule forces to the full suite short-circuits before the cache is read, so its contexts block records "consulted": false rather than a baseline of null — an escalated run executed every test and was never exposed to a narrow selection, and counting it as one that ran on the static closure alone is what used to inflate the exposure reading below.

Line numbers are read on the baseline side of git diff -U0 <baseline-sha>, restricted to the change set's own files, because the database is keyed by line numbers as they were in the baseline.

When there is no usable baseline

A run that finds no usable baseline narrows on the static closure alone, which the backtest below measures as a real blind spot — so it cannot simply carry on as if the closure were sound. How often that happens is worth stating carefully, because the first reading of it was wrong: just selection-health used to count every escalated run as one without a baseline, which made absence look like half the lane. Escalated runs never consult the cache and run every test anyway. Over the same store measured by consulted runs only, a baseline was present in 21 of 23; the 21 remaining scoped runs escalated before contexts could matter.

So absence is uncommon on a host that fetches or records baselines — but it is not rare where it counts. It is the standing condition of a workspace that has been idle past the CI artifact's 14-day retention, one that is offline, or a host that has never fetched, and there absence is persistent rather than occasional. Escalating on it would be sound and is now known to be affordable at this frequency; the closure walks one hop deeper instead because a measured 91% of the blind spot comes back for roughly double the selected files, against 3,650 worker-seconds for a full run. That records no-baseline-depth-boost, which appears in the manifest, in just check's scoped summary line, and in just selection-health's rule histogram. The manifest's effective_depth is the depth actually walked, configured depth plus whatever the rename/delete and no-baseline compensations bought; depth stays the configured one.

Measured with just selection-backtest --limit 150 --include-descendant-baseline --baseline 96183d71b at 4651ed199 over 63 commits with usable ground truth (3 faithful baseline-ancestor replays, 60 approximate baseline-descendant ones), closure-only:

depth mean recall p10 recall worst recall blind-spot commits missed test files median selection
2 (before) 96.0% 85.3% 23.5% 13 / 63 116 6.4%
3 (with the boost) 99.2% 100.0% 81.3% 5 / 63 11 8.8%

The extra hop costs roughly double the selected files (src/sase/agent_lanes.py: 110 → 255 of 2,329, 1,117 → 2,514 tests, 57s → 164s serial on athena) and raises the replayed escalation rate from 23/63 to 28/63 — historical whole-commit diffs, well above what a working-tree change selects. It buys back 91% of the measured blind spot. On the sharpest known shape, src/sase/ace/tui/_app_layout.py — widely executed but shallowly imported — it lifts recall from 24.2% to 53.8% (69 missed of 91 down to 42) at 14.2% of the suite, still under the escalation ratio. Directory-mirror expansion was measured as the alternative for that shape and rejected: tests/ace/tui/** is 831 files, 35.7% of the suite, so mirroring escalates to the full suite rather than staying scoped.

The contract set

Some tests audit the repository as a whole rather than one module: config-schema conformance, generated-file drift, terminology guards, tool-script contracts. No import edge connects them to the code they police, so the closure would never select them. They are marked @pytest.mark.contract and added to every scoped selection unconditionally.

tests/contract_manifest.txt is a generated projection of that marker, not a hand-maintained list — the selector reads the committed file so it does not have to collect the suite first. To add or remove a test file from the set, change the marker on the test module and regenerate:

just refresh-contract-manifest   # rewrite tests/contract_manifest.txt from -m contract

tests/test_contract_manifest.py fails when the committed manifest disagrees with the marker, so a forgotten refresh surfaces as a test failure rather than a silently stale selection. The same module carries a budget guard bounding the size of the set: every agent pays for it on every just check, so growth has to be deliberate. The guard asserts a manifest-entry cap calibrated from the current measured serial cost instead of timing a nested contract run; timing proved too load-sensitive to use as a correctness oracle under real xdist contention.

Both guards live outside the contract set on purpose — regenerating and re-budgeting the whole set from inside that same set would charge every just check for it twice — so they run only in the exhaustive lane (just test, just check-full, CI). Marking a test and forgetting the refresh therefore survives a just check: CI is the backstop. just check already escalates when tests/contract_manifest.txt itself changes.

Once you do regenerate, the manifest change broadens the next selection by itself. tests/contract_manifest.txt belongs to the selection-tooling broadening rule, so the just check that lands a new contract test escalates to the full suite — that escalation is the manifest edit, not the new test.

The contract set is also the floor. A change set that contributes no import-graph seeds at all — a docs-only edit, an sdd/** change, a .github/** workflow tweak — records the contract-set-only rule and runs exactly the contract tests. That rule does not escalate: running the whole suite for a Markdown edit would be the heuristic failing in the expensive direction.

Expect selections to grow. Over the 2026-08-06 baseline, 1,237 of the ~2,400 measured src/ files have at least one line whose per-test contexts (40 tests or fewer) include a test the depth-2 closure never selects. The sharpest case is src/sase/ace/tui/widgets/_file_completion_refresh.py, where the closure selects zero test files and contexts select 40 — a change there would previously have been checked by nothing but the contract set. In the other direction, a line in a widely-executed module really is executed by thousands of tests, and contexts say so, which will push some selections over the escalation ratio and into the governed full lane. Both directions are the heuristic being corrected rather than a malfunction; watch just selection-health for what it does to the escalation rate and the false-negative count.

SASE places a pytest safety boundary around its telemetry mutations and common axe state/log writers when they target the OS account's real ~/.sase tree. Telemetry flushes and deletions fail with an actionable error; guarded best-effort daemon writes are suppressed and warn once per target and category; and axe start, stop, and restart requests are refused unless their test-only override is set. The pytest harness also publishes SASE_PYTEST_SANDBOX_DIR; while that marker is present, bead-store writes through the Python mutation facade or Rust CLI fast path are refused unless the target store is at or below the sandbox root. SASE_ALLOW_UNSANDBOXED_BEAD_WRITES=1 is the deliberate test-only escape hatch for a genuine exception. SASE preserves pytest's isolation marker when it starts runner and daemon subprocesses. These guards are not a substitute for isolation: tests that exercise persistence should point SASE_HOME at a per-test temporary directory and create bead stores under tmp_path or another path inside the published sandbox. The default fixture also sets SASE_DISABLE_PLUGIN_CONFIG=1 so tests assert bundled defaults rather than whichever plugins happen to be installed. Request the real_plugin_config fixture when the subject is production merge of a plugin sase_config layer. Run just test-bead-store-soak when changing bead resolution or mutation paths; it runs the default suite and verifies the legacy production plans sidecar's beads/issues.jsonl digest, bead-state git status, and git HEAD are unchanged. The current guard still targets SASE_SDD_PLANS_DIR/beads and never resolves the dedicated beads role. On a cleaned schema-3 project it exits with a missing-file error instead of running the suite; if a legacy beads/ copy remains under --plans, the helper guards that stale copy rather than the active --beads store. If an older test run already polluted the telemetry store, preview the exact-label cleanup with sase telemetry cleanup-test-data --dry-run before deciding whether to rerun it with --yes.

just test, just test-slow, just test-visual, and just test-cov share a host-global pytest-xdist worker-token pool with every other checkout owned by the same UID. An automatic run waits until it can lease a small floor, then greedily grows to its per-run ceiling using whatever capacity is currently free. On the standard development host, a solo run can receive 28 workers from the 32-token pool while leaving four tokens for another run; concurrent runs scale down to their actual grants instead of each independently oversubscribing the host. The granted count is the value passed to pytest -n.

The default host budget reserves max(1, cpu_count // 8) CPUs and 8 GiB of available memory, allows 950 MiB per worker, and never exceeds 32 tokens (the prior safe aggregate ceiling). The CPU reserve is proportional rather than a flat count, so a small host (e.g. a 4-vCPU CI runner) still gets real parallelism instead of collapsing to a single worker. The memory allowance was calibrated from live worker RSS sampled across concurrent sibling workspaces, which ranges from 0.74 to 0.85 GiB; 950 MiB keeps headroom over the top of that range. Missing memory information falls back to a conservative four-token limit, and small hosts clamp to at least one token. These capacity safeguards are independent of xdist scheduling and individual test cost.

The runner defaults to pytest-xdist's worksteal scheduler. Workers begin with evenly divided queues and can reclaim pending tests from a worker with a long queue, avoiding the idle-worker tail caused by keeping an entire heavy test file on one worker. The fallback is a one-variable change:

SASE_PYTEST_DIST=loadfile just test

SASE_PYTEST_DIST accepts only worksteal and loadfile; unsupported values fail with a pytest usage error before the runner leases worker tokens. Inline-snapshot update and review modes remain serial and omit both -n and --dist regardless of this setting. Test selectors and other pytest options continue to pass through normally.

A post-change comparison on 2026-07-20 used the same 19,883-item fast-suite selection, refreshed dependencies, and an exact governed grant of 28 workers. Aggregate CPU is reported as the mean utilized cores divided by the 28-worker grant; the tail is wall time from the first 99% progress report through completion.

Scheduler Pytest time Wall time Grant utilization 99%-finish tail
loadfile 109.68s 111.93s 59.1% 41s
worksteal 102.34s 104.66s 61.8% 39s

worksteal reduced wall time by 7.27s (6.5%) while running the same assertions. The slowest calls remained the two tests in test_agents_zoom_panel_search.py (roughly 16-20s each), so the improvement reflects better pending-work distribution rather than removed test cost. Three complete worksteal runs at governed grants of 11, 16, and 28 workers passed while auditing for within-file order and shared-state assumptions.

Final combined-suite verification

The completed optimization was measured on athena on 2026-07-20 with the current 19,921-item fast-suite selection. A crash-safe reservation held the measured 29-token host pool across three consecutive samples so unrelated queued suites could not enter between runs; each just test used 25 workers, the automatic ceiling for that capacity after reserving the four-worker floor. Aggregate CPU is GNU time's mean utilized cores, and grant utilization divides it by 25.

Sample Pytest time Recipe wall Workers Aggregate CPU Grant utilization
1 90.71s 93.14s 25 1809% 72.4%
2 90.84s 93.08s 25 1776% 71.0%
3 89.63s 92.05s 25 1803% 72.1%
Mean 90.39s 92.76s 25 1796% 71.8%

Against the pre-optimization 4:04 recipe / 194s pytest / 14-worker / ~780% CPU baseline, the mean recipe is 2.63x faster and the pytest segment is 2.15x faster. Non-pytest recipe overhead fell from roughly 50s to 2.37s. The selection grew from 19,744 to 19,921 items while the work landed; no existing test was removed, skipped, or moved out of the fast lane.

Coverage parity used the same 19,921-item selection with just test-cov: 19,915 tests passed, 7 were skipped, total branch coverage was 80.07%, and the unchanged 50% gate passed. just test-cov shares just test's marker selection, which at the time of this measurement still included sase's TUI PNG visual regression tests. Both recipes exclude those tests today; see Visual Snapshot Workflow.

Sustained real-host demand also exercised the pool while these measurements were prepared. With memory sizing the active budget at 20 tokens, three full suites progressed simultaneously with grants of 12, 4, and 4 workers. Their sum never exceeded 20; available memory stayed healthy and swap remained at 2.3 GiB throughout the observation. The process-level regression in tests/test_suite_gate_scaled_integration.py makes the same guarantees deterministic in a temporary three-token pool: three one-worker suites reach test execution together, a fourth waits, killing one holder admits the waiter, and active grants remain exactly bounded before and after the handoff.

Set SASE_PYTEST_WORKERS=<N> to request exactly that many governed workers; the request must fit the shared capacity. Direct parallel pytest -n ... controllers use the same pool and lease their resolved numeric, auto, or logical worker count exactly. Lock descriptors survive the runner's exec and are released by the kernel even after SIGKILL. Nested pytest processes inherit the disabled marker so they cannot deadlock on the parent's tokens.

For deliberate diagnostics, SASE_TEST_GATE_DISABLED=1 bypasses accounting: the run takes no tokens and never queues. It is still clamped to the host budget and prints one line saying so, because the pool cannot see it and every other run's budget assumes it is absent — an unaccounted 64-worker controller against a 32-token pool once drove this host to a load average of 97.6 with 25 GiB in swap. A benchmark that genuinely needs to run wider raises SASE_TEST_GATE_SLOTS instead, which enlarges the pool where concurrent runs can see it. A run whose exemption is corroborated by a real ancestor lease (SASE_TEST_GATE_GOVERNED=1, or an xdist worker) is unaffected: its width was already paid for, so it is granted untouched.

SASE_TEST_GATE_SLOTS overrides host-wide token capacity, SASE_TEST_GATE_TIMEOUT controls bounded admission waits, and SASE_TEST_GATE_DIR selects the shared pool directory. A live holder is also bounded: SASE_TEST_GATE_STALE (default 30 minutes without a progress heartbeat) and SASE_TEST_GATE_MAX_HOLD (default 4 hours, even while heartbeats continue) reclaim a wedged grant so one sleeping tools/run_pytest cannot keep a third of the host pool overnight. SASE_PYTEST_WORKER_FLOOR and SASE_PYTEST_WORKER_CEILING tune automatic grants; invalid or inconsistent values fail before pytest starts. See Configuration for the complete contract.

Test selectors are normalized from the directory where just was invoked, so this works the same from the repository root or a subdirectory:

just test tests/main/test_parser.py::test_example

just fix, just fmt, Ruff formatting/checks, generated model-alias docs, and keep-sorted use a narrow formatter environment at .venv-format/ by default. That environment installs only the format-tools dependency group plus the repo-local Prettier and keep-sorted bootstraps, so local formatting does not rebuild the Rust extension, install plugins, or depend on the application .venv/. Override the formatter environment with SASE_FORMAT_VENV_DIR or Just's format_venv_dir variable.

just lint and just fix-keep-sorted bootstrap a project-local keep-sorted executable into the formatter environment from PATH, or by running go install github.com/google/keep-sorted@v0.8.0 when Go is available. If neither keep-sorted nor Go is installed, those recipes fail with a setup error before linting YAML keep-sorted blocks. Runtime, test, install, and mypy recipes still use .venv/ and the full _setup path because they need the installed application environment and Rust binding validation.

Default test runs select not slow and not visual, so sase's TUI PNG snapshot regression tests do not run in just test, just test-cov, or just test-scoped. just fix-tui-screenshots is the canonical visual execution path; just test-visual is the check-only alias. Both install the optional PNG rasterizer dependencies when they are missing. The real-PTY smoke tests carry both terminal_smoke and slow, so that same expression excludes them too — terminal_smoke selects them, it does not deselect them. Direct pytest runs inherit the identical default expression from pyproject.toml unless you pass your own -m selector.

Use just test-terminal-smoke only when you need to verify sase's TUI startup path through a real PTY. It installs pexpect and pyte, runs the optional terminal_smoke marker, and stays out of default tests and CI until that path has proved stable. The recipe uses the shared pytest runner's private disk-backed temp root and leak guard, but it is always serial and never leases xdist worker tokens; SASE_PYTEST_DIST is therefore ignored. Set SASE_PYTEST_TMPDIR to override its scratch root while diagnosing temp-path behavior.

The Master Gate (SASE_TEST_SHARD)

.github/workflows/master-gate.yml is a second, additive gate that runs on every push to master, grouped by master-gate-${{ github.sha }} with cancel-in-progress: false so every SHA gets its own bounded run that is never cancelled by a newer push — unlike ci.yml's per-ref concurrency group, which lets a newer push replace an older one's pending slot. It runs the complete non-visual fast suite, split into eight deterministic, balanced shards (SASE_TEST_SHARD=<1-based index>/<1-based count>, e.g. 3/8) so no single job has to carry the whole suite's wall clock, and reuses a SHA-keyed sase-core wheel cache so a sase-core revision that has not moved costs one cache restore instead of a full Rust rebuild.

tests/_test_shards.py does the splitting. It walks tests/**/test_*.py on disk (not git ls-files, so an uncommitted new test file is still in scope), estimates each file's cost from the committed tests/shard_timings.json, and assigns files to shards with longest-processing-time-first: the files are sorted once by descending cost estimate (ties broken by a SHA-256 digest of the path, for a fully deterministic order), then each one is dropped into whichever shard is lightest so far. Discovery is always exhaustive — an unrecognized or stale timing table can only unbalance the shards, never drop a file from all of them — so SASE_TEST_SHARD is deliberately supported only in just test's fast mode and refuses an explicit test selector alongside it: sharding answers "which slice of everything," not "which subset."

tests/shard_timings.json is generated from this host's local per-test-file duration recordings (the same store Per-test-file timings describes), because a fresh CI runner has no local history of its own. tools/refresh_shard_timings retains the 800 slowest measured files individually (rounded to 0.1s) and folds everything else into one default_duration — the mean of exactly the files it does not name, not the mean of the slow files it does. Run it after a just test or just check-full that already recorded fresh timings:

just refresh-shard-timings                # write tests/shard_timings.json
just refresh-shard-timings --check        # verify it is not stale, without writing
just refresh-shard-timings --print-plan 8 # preview an 8-shard split

A file the table has never seen — new, renamed, or simply outside the retained 800 — still runs; it just costs default_duration instead of a measured number, so the worst a stale table does is uneven shards, never a skipped test. tests/test_test_shards.py polices staleness directly: it fails if fewer than 90% of the committed table's retained files still exist, if the discovered file count has drifted more than 20% from the table's recorded count, or if the master gate's eight shards built from the committed table are unbalanced by more than 10% of their mean — every failure points back at just refresh-shard-timings.

The table also has a CI freshness path so it cannot silently decay after those thresholds. Full CI's 3.14 just test leg publishes the folded table as the shard-timings artifact; a weekly shard-timings-ratchet.yml workflow copies that artifact into tests/shard_timings.json when the proposed table would change the gate's file split, or when the committed generated_at is older than 14 days, and opens a PR. just refresh-shard-timings --from-payload PATH --check --assignment is the same comparison the ratchet runs.

The master gate's lint job is kept byte-for-byte identical to ci.yml's own lint job steps (a contract test in tests/test_github_actions_ci_master_gate.py polices the equality), so this gate's lint signal cannot drift from what PR CI already promises. It does not run the diff-scoped lane, coverage, slow, or visual lanes — those stay on ci.yml's own jobs and schedule. Cost attribution and coverage-contexts moved further out still, onto telemetry.yml's own schedule.

Reproducing Timing Flakes (just test-contention)

The default lane's timing flakes are a class, not a list of nodes: individually rare, collectively frequent, and historically only reproducible under accidental host load. just test-contention makes them reproducible on demand, the same way just test-visual-contention already does for PNG convergence: taskset pins a 26-worker pool to two CPUs (13x oversubscription), the selection runs N times, and the run ends with a per-node tally naming each failing node, how many repeats it failed in, and which ones.

just test-contention -- tests/ace/tui/util/test_stall_watchdog.py   # restrict the soak
SASE_CONTENTION_REPEAT=6 just test-contention -- tests/test_bead    # soak harder

Override the pinned CPU list, the worker count, and the repeat count with SASE_CONTENTION_CPUS, SASE_CONTENTION_WORKERS, and SASE_CONTENTION_REPEAT (defaults 0,1, 26, and 3). A full-suite repeat is far too slow to iterate against, so pass paths or node IDs; the tally is what turns "it went green once" into a before/after measurement a fix can be falsified by.

Per-repeat failure records land in .pytest_cache/sase-contention/repeat-NN.json, so a finished soak can be re-read without re-running it.

This lane is an opt-in diagnostic and is deliberately kept out of every governed path: it takes no suite-gate lease, writes nothing to the durable selection-health store, and is unreachable from just check and just check-full. A deliberately starved run is not evidence about what a scoped run should have selected. It also starves the host on purpose, so other agents' runs on the same machine slow down while it runs.

Selection Health

just test-scoped selects tests from the change set with a depth-bounded reverse walk of the import graph. That selection is a heuristic, so its cost and its mistakes are both measured rather than assumed.

Every scoped run copies its selection manifest, and every full-lane run (just test, just test-cov, just test-cost) copies the node IDs it saw fail, into a durable host-local store at ${SASE_HOME:-~/.sase}/test-selection/<project-key>/. The store is shared by every numbered workspace of the project, so the report reads one project-wide sample rather than one workspace's, and records older than 30 days are pruned on write. Sharing the store is not the same as correlating across it: records carry the workspace and change set that the false-negative rule below needs precisely so that one workspace's flake is never charged to another workspace's selection.

just selection-health          # readable report
just selection-health --json   # the same numbers, machine-readable

The report covers how many scoped runs ran, how often they escalated to the governed full lane, median and p90 selection size, scoped duration percentiles with the middle gear's width mix behind them, worker-seconds of host demand avoided (charged at the leased width, so a 100s run at four workers costs 400), which broadening rules fired, and — the number that decides whether the fast lane is trustworthy — the false negatives: tests that failed in a full run after a scoped run over the same change excluded them. The target is zero of a sample that already excludes known flakes (see below). A non-zero count means the heuristic itself is unsound as tuned; the response is to raise SASE_TEST_SELECTION_DEPTH to 3, or mark the missed tests @pytest.mark.contract and run just refresh-contract-manifest, and then re-measure — not to explain the failures away.

The report also names the lane's own worst behaviour instead of letting the median hide it: alongside p75/p90/max scoped duration it prints "scoped runs slower than the full lane (FULL_LANE_WALL_SECONDS)", with each offending run's selected-file count and the rules that produced it, so a latency regression like the one budget was built to fix is visible in the project's own health metric rather than only in one-off timed measurements. Escalated runs are called out separately as "cost not measured" rather than folded into the percentiles at their recorded duration: 0.0 — that zero is a placeholder for a run handed off before the runner could time it, not a real duration, and counting it as fast would silently hide exactly the regression this counter exists to show.

"The same change" is what makes that number mean anything, and the report states the rule on every run: a scoped run is charged with a full-run failure only when both records name the same workspace, the scoped run's HEAD is an ancestor of the full run's, and the full run's change set covers the scoped run's. Ancestry alone is not enough — sibling workspaces normally sit on the same master HEAD, so is_ancestor(head, head) is trivially true and every workspace's flakes would be charged to every other workspace's selection.

Read the count together with the two lines under it. Records written before health schema 2 carry no workspace or change set, cannot satisfy the rule, and are excluded from correlation; the report says how many there are, so a zero is read as zero-of-a-known-sample rather than mistaken for a clean one.

A failure that clears all of the above can still be a known flake: reproducible_flake_nodeids (tests/_test_selection_health.py) looks at every full run that saw the same node fail, and calls it reproducible when those failures span unrelated change sets and an independent full run between them passed the node. That second requirement keeps a deterministic master break, which fails everywhere until its fix lands, from being counted as host-load flakiness just because multiple workspaces hit the same bad commit range. Matches on a reproducible node are moved out of the false-negative count and into a separate flake-suppressed line, counted and listed exactly like the false negatives are, never silently dropped. A single occurrence is never enough evidence on its own and stays a false negative until it recurs. This needs no hand-maintained list of node IDs — the real store already showed failures reproducing on nodes no bead had enumerated yet, including one caused by a stale sase_core_rs build rather than test-isolation timing, so a fixed list would already have missed it. (A missed test still charged by exactly one scoped selection's change set, rather than reproducing across full runs, gets the older, softer hint instead — "matched across unrelated changes; suspect a flake before a miss" — since that alone is not enough evidence to suppress.)

The coverage contexts block reports baseline availability over the runs that consulted the cache, not over every scoped run, and states separately how many escalated before contexts could matter. Those two denominators differ by roughly the escalation rate — on a store where half the runs escalate, counting them as baseline-less made a lane with two genuinely closure-only runs read as twenty-three of them.

Use tools/select_tests --explain to see why an individual test was or was not selected. Set SASE_TEST_SELECTION_HEALTH_DISABLED=1 to skip recording entirely, and SASE_TEST_SELECTION_HEALTH_DIR to point the store somewhere else.

The Flake-Baseline Gate

The reproducible-flake set is not only reported — it is gated. After the full test lane, just check-full runs just selection-health --fail-on-new-flake, which compares the currently reproducible node IDs against the committed baseline at tests/reproducible_flake_baseline.txt and exits non-zero on any node that is not already listed there. The failure names each new node and ends with Additions require a filed bead; fix or file the node before landing. — so the response to a red gate is to fix the test or file a flake task bead and add the node with a comment naming that bead, never to add the node silently. After that gate succeeds, just check-full runs just fix-tui-screenshots so local exhaustive verification can refresh screenshot goldens.

The gate deliberately judges a narrower sample than the report:

  • Only full-lane records that carry a change set (health schema 2+) and at most five failures are eligible, so one broken suite run cannot promote every node it touched into flake debt.
  • It needs at least two eligible records. With fewer, it prints not enough full-lane records to judge and passes.
  • The baseline file's # effective-after: <UTC timestamp> line discards evidence recorded at or before that instant, which is how the list is reset after a suite-wide fix.
  • A # fixed-at: <UTC timestamp> <node id> line retires only that node's pre-fix evidence. A later failure of the same node is ordinary live evidence again.
  • Node IDs that no longer name a collectible test (renamed or deleted) are reported as stale rather than counted, so a rename cannot manufacture permanent pressure to bump the cutoff.

Entries in the baseline are debt to remove, not suppressions to grow. Run just selection-health (without the flag) for the readable report behind the verdict.

Selection Backtest

just selection-health's false-negative count can only grow when a full run happens in the same workspace as an earlier scoped run over a subset change. In ephemeral workspaces that combination essentially only occurs at landing, so the correlatable sample grows about as fast as epics land. just selection-backtest answers the same question from history instead, today.

just selection-backtest                                  # replay the last 50 commits
just selection-backtest --limit 150                       # a longer window
just selection-backtest --json                            # the same numbers, machine-readable
just selection-backtest --execute --execute-limit 1       # actually run the missed tests

For each replayed commit the harness checks the commit out into its own throwaway detached worktree (never the invoking checkout), takes the commit's own diff against its parent as the change set, rebuilds the import graph as of that commit, and computes the selection the scoped lane would have produced. Ground truth for the same change set comes from the cached coverage baseline: the test files coverage recorded as executing the lines that commit touched. Recall is the share of that ground truth the selection contained.

Recall is reported twice. closure-only runs with the contexts cache forced absent and is what a workspace with no cached baseline actually gets. closure+contexts is 1.0 by construction — the selector unions in the very same coverage query the ground truth comes from — so it is not independent corroboration. The gap between the two arms is the exposure, and it is the number a compensating action for a missing baseline has to be tuned against.

Three limits bound what a reading proves, and the report states each of them rather than burying them:

  • Ground truth needs a usable baseline. By default only commits the baseline is an ancestor of are replayed. Since baselines arrive as a CI artifact on master pushes, that is a small window. --include-descendant-baseline also replays commits the baseline sits ahead of; ground truth for those is widened by every later change to the same file, so recall reads pessimistically, and the report counts the two directions separately.
  • The replay is conservative. core-identity-changed cannot fire historically — the venv a commit was tested against is gone — so runs that escalated in reality may replay as narrow selections. The harness under-reports recall.
  • Recall is a proxy. A missed test file is a true false negative only if it would have failed. --execute checks that for the worst few blind spots by running the missed files at their commit. It is opt-in, slow, and deliberately absent from just check and just check-full (tests/test_justfile_lint.py pins that).

Measured on 2026-08-06 at 6b0976bcb, over --limit 150 --include-descendant-baseline against the 96183d71b baseline — 65 commits with usable ground truth (1 faithful, 64 reverse-direction), 85 skipped and itemised:

arm median recall mean p10 worst commits with a blind spot missed test files
closure-only 100.0% 96.2% 86.7% 23.5% 13 / 65 118
closure+contexts 100.0% 100% 100% 100% 0 / 65 0

25 of the 65 reached perfect recall by escalating rather than by selecting well. The worst case — 6719992521ad, feat(sidecars): surface publication queue observability — recalled 23.5%, missing 75 of 98 covering test files. Median selection size was 6.4% of the suite (p90 11.9%), so the closure-only arm is not paying for its misses with breadth.

Note what the skip counts say about the sample: of the 150 commits examined, 46 changed no src/**.py at all and 36 touched no file with a baseline-side line to query. A recall figure here is a figure over commits that change already-covered production code, not over all commits.

Visual Snapshot Workflow

sase's TUI visual tests live under tests/ace/tui/visual/ and tests/pager/visual/ and compare deterministic Textual screenshots against committed PNG goldens. ACE goldens live in tests/ace/tui/visual/snapshots/png/; pager goldens live in tests/pager/visual/snapshots/png/. The renderer stack is exact-pinned in the visual optional-dependency group in pyproject.toml, and tests/ace/tui/visual/renderer_env.json records those package versions plus hashes of the bundled fonts. A session-scoped fixture checks that fingerprint before any snapshot runs, so a skewed environment fails once with an installation or upgrade instruction instead of producing a wall of misleading pixel diffs.

The visual fixtures also pin the process environment that affects rendering: TERM=xterm-256color and COLORTERM=truecolor select Rich's truecolor path, FORCE_COLOR and NO_COLOR are removed, and TZ=UTC is applied with the process timezone cache refreshed. Neither a contributor's terminal settings, local timezone, nor CI's process environment participates in the golden corpus.

just fix-tui-screenshots is the canonical maintenance command. It captures both visual trees, compares candidates with exact pixel equality, and on Linux applies every golden it can prove. Update mode salvages per node and per golden:

  • Capture retry. If the first pass leaves no usable inventory at all, the whole capture pass is retried once (capture-retry/); a second empty result fails the run.
  • Node recovery. Visual nodes that failed, errored, or were collected but never accounted for (for example, stranded on a lost worker) are rerun up to two more times (recover-1/, recover-2/), serially once 25 or fewer remain. Only nodes that pass are trusted; the rest are skipped and their goldens left untouched.
  • Concurrent edits. Once capture and node recovery finish, the golden trees are compared with the run-start snapshot; a golden that changed on disk in that window is skipped rather than overwritten. This check runs once, before verification: an edit made during verification or apply is not detected and can be overwritten, so do not edit goldens while a run is in progress.
  • Agreement voting. Each remaining created or updated candidate is recaptured by its owning test in up to three verification passes (verify/, then serial verify-2/ and verify-3/) and is applied once two of its captures, counting the original, agree byte-for-byte; a candidate whose first recapture matches needs only one pass. A golden whose captures never agree, or whose recapture comes from a different owner, is skipped.

Everything skipped is listed in a WARNING block under status partial, and the run still exits 0. In update mode, a selection that matches no visual tests exits 0 with a warning instead of an error; check mode treats it as an execution failure (exit 3). -n N / --numprocesses N (with or without a leading --) is translated to SASE_PYTEST_WORKERS=N for the governed runner rather than reaching pytest (-n auto and -n logical fall back to the governed default); -n inside PYTEST_ADDOPTS is a usage error. Pass --check to inventory the same way without writing goldens; check mode does no salvage, stays strict, and exits 1 on required drift. Arguments after -- are pytest selectors (paths, node IDs, -k). Targeted runs apply only captured changes and never prune unvisited files. A requested full run applies creates and updates, but stale removal needs complete evidence — every collected node trusted (skipped and xfailed nodes excepted), no capture protocol errors, and trusted captures from each root that has goldens; otherwise pruning is skipped with a warning, the run reports no stale entries, and the status is partial. When another run in the same checkout holds the maintenance lock, the runner prints a waiting notice and waits (bounded, 2 hours) instead of refusing at once; timing out exits 2.

just fix-tui-screenshots
just fix-tui-screenshots --check
just fix-tui-screenshots -- tests/ace/tui/visual/test_ace_png_snapshots.py -k example

just test-visual is a supported check-only alias of --check. just update-visual-snapshots is a supported update alias of the same runner. just test-visual-contention remains a check-only diagnostic that uses the governed visual pytest runner directly. just check, just lint, just fix, and the per-SHA Master Gate do not run screenshot maintenance.

Local just check-full runs the update form after the other exhaustive gates succeed and prints the compact report (scope, status, counts, report path) outside tools/run_silent. That stage can modify committed goldens. It exits 0 with status partial when goldens are left untouched behind warnings; read the WARNING block — those goldens are not known to be current. Direct check-full in CI refuses at the update stage; repository CI uses just fix-tui-screenshots --check instead. Update mode also refuses when GITHUB_ACTIONS is set, when CI is set outside a SASE agent workspace, sase monitor command, or SASE proc, off Linux, or when the renderer fingerprint is skewed. SASE agent processes export CI=true for pytest/tooling; that flag alone does not block local golden updates. Detached monitor commands and detached ToolRun procs (an agent's escalated or -d sase tool run) do not inherit SASE_AGENT*, but they set SASE_MONITOR_ID or SASE_PROC_ID, which is treated the same way. --sase-update-visual-snapshots is retired; pytest rejects it with the replacement command.

Every update run (clean, applied, or partial) and check-mode drift retain a reviewable report under a unique run directory in .pytest_cache/sase-visual/runs/, alongside the logs and candidates of every pass (capture.log, capture-retry.log, recover-N.log, verify*.log). .pytest_cache/sase-visual/latest-report.json points at the current run only. At the end of every run, while still holding the maintenance lock, old run directories are pruned: the run latest-report.json points at, the current run, any run with an unfinished (planned or applying) or unreadable apply journal, the 10 most recent runs, and any run with a file modified in the last 24 hours survive, and the rest are deleted. The manifest records warnings, skipped (each with a reason such as test_failed, unstable, owner_mismatch, or concurrent_edit, plus evidence paths), attempts, and pruning_skipped_reason, and the HTML report and summary.md end with a "Not updated" section listing the same skips. Inspect every creation and removal, then each update group (representative plus members), then the "Not updated" list — generation is not approval, and a skipped golden is not known to be current. After an interrupted apply, the next update invocation may restore the recorded baseline when hashes still match; if the journal conflicts with the current goldens, update mode refuses with exit 2. A check invocation never performs recovery writes and fails with exit 3 while an unfinished journal exists.

just may normalize a non-zero child code to 1. Automation that needs the distinction between drift (1), usage/environment refusal (2: bad arguments, a pytest usage error, CI/platform/renderer refusal, an update-mode journal conflict, or a lock-wait timeout), and execution/application failure (3: no usable inventory after the capture retry, an apply failure after rollback, an interrupt, or, in check mode, a failed, empty, or incomplete capture or an unfinished journal) should read the run manifest or invoke tools/fix_tui_screenshots directly (--help prints the same contract).

Committed goldens are canonical to the pinned renderer. Rasterization goes through resvg (resvg_py==0.3.3), a pure-Rust SVG renderer that carries its own font database restricted to the bundled fonts in src/sase/ace/tui/fonts/ with skip_system_fonts=True. No host font-config or graphics stack participates, so rendering is stable and host-font-independent on the canonical Linux x86_64 platform. Fira Code is named for every generic family, so it wins every glyph it carries; DejaVu Sans is bundled purely as the fallback resvg reaches for on a codepoint Fira Code lacks. Without it, symbol marks such as the notification tab icons would rasterize as missing-glyph boxes in every golden while rendering correctly in a real terminal, and no reviewer could tell the two apart by eye. tests/ace/tui/visual/test_tab_icon_glyphs.py makes that check mechanical: it fails if the bundled fonts stop covering an icon sase's TUI can pick without configuration. PNG comparison is byte-exact by default locally and in every visual-bearing CI lane; together with the fixture-level terminal and timezone pins, a mismatch is a real rendering change or an unpinned environment defect to investigate.

Rasterization can still differ by a small, bounded amount on macOS arm64. The tolerance environment variables remain available only as explicit escape hatches for local iteration and renderer investigations. For the known macOS drift, use:

SASE_VISUAL_PNG_MAX_DIFF_RATIO=0.01 \
SASE_VISUAL_PNG_MATERIAL_DIFF_THRESHOLD=8 \
SASE_VISUAL_PNG_MAX_MATERIAL_DIFF_PIXELS=0 \
just test-visual

The ratio caps the changed image area. The material threshold measures the maximum visible channel distance after alpha-aware compositing over black and white, and the material-pixel cap still rejects any change above that threshold. These overrides never update or implicitly accept a golden, and they do not bypass the Linux-only regeneration gate. Per-assertion equivalents are max_diff_pixels, max_diff_ratio, max_material_diff_pixels, and material_diff_threshold.

Mismatch assertions, summary.txt, and failure.json report material_diff_pixels, material_diff_ratio, and material_diff_threshold alongside the active area and material limits. Inspect those fields to distinguish broad, low-amplitude renderer drift from a small material UI change before using any override.

One accepted fidelity caveat: Fira Code ships no italic face and resvg does not synthesize oblique, so font-style: italic renders upright. This is uniform across every screen and host. Restoring visible italics would mean switching the bundled font family, taken as a separate follow-up if it becomes necessary.

A second one no longer applies: emoji-presentation codepoints were uncovered because a deterministic rasterizer cannot use a color-emoji font, but the monochrome Noto Emoji static outline face (bundled as NotoEmoji-Regular.ttf) rasterizes them the same way the other bundled fonts do. tests/ace/tui/visual/test_emoji_glyphs.py audits every emoji codepoint src/sase actually uses the same way test_tab_icon_glyphs.py audits tab icons.

Intentional Renderer Upgrades

The pinned versions and font bytes define the golden corpus. Upgrade Textual, Rich, resvg, a syntax grammar, Pillow, or another package in that stack as one reviewed change:

  1. Update the exact pins in the visual optional-dependency group in pyproject.toml.
  2. Run uv lock, then just install-venv-visual so the working environment matches the new pins.
  3. Refresh the matching package versions in tests/ace/tui/visual/renderer_env.json. If bundled fonts changed, update their SHA-256 hashes too; the Python and platform fields are diagnostic only.
  4. On Linux, run just fix-tui-screenshots, inspect the report, then run just fix-tui-screenshots --check and require an unchanged golden tree.
  5. Review the complete PNG diff for unexpected content or layout changes and commit the pins, uv.lock, fingerprint, and regenerated goldens together.

Non-Linux contributors should use CI as the canonical renderer. Push the branch, let the Linux visual-test job produce ace-visual-artifacts, and download that artifact from the Actions run. The uploaded run directory contains the maintenance report, logs, and partial captures. Review those instead of accepting goldens from a failed CI execution. The same fingerprint checks still require pins, lockfile, and manifest to agree before CI will render the replacement corpus.

CI Visual Lanes

The default lane (just test, just test-cov, and every leg of the Python matrix) excludes visual tests. The dedicated Linux Python 3.12 visual-test job is the check-only visual execution: it runs just fix-tui-screenshots --check and uploads the current run's report plus raw captures. It never writes goldens, never relaxes equality, and never implies that an execution failure is merely unaccepted corpus drift. This keeps one broad lane plus one diagnostic lane authoritative for snapshots while preventing a future Python-specific rendering change from reddening the whole matrix.

Visual Failure Report

tools/render_visual_snapshot_failure_report still consumes legacy failure.json sidecars from direct diagnostic comparisons. Screenshot maintenance feeds it a versioned run manifest instead. Each run writes self-contained HTML, Markdown summary, JSONL change records, and images under that run's report/ directory:

  • visual-failure-report.html - self-contained HTML with PNG/SVG embedded as data URIs.
  • summary.md - compact table for $GITHUB_STEP_SUMMARY.
  • annotations.sh - escaped ::error file=...,line=... workflow commands.
  • manifest.jsonl - aggregate of every loaded change or failure record.

Created images show that there was no baseline; stale removals show the previous image and lack of a producing assertion. Neither is labeled as an ordinary pixel diff.

Run it locally against a maintenance manifest with tools/render_visual_snapshot_failure_report --manifest <manifest.json> --repo <owner/repo> --sha <commit> or against legacy sidecars with --artifact-root. The script is safe to run when there are no failures; it exits 0.

In GitHub Actions the visual-test job reads .pytest_cache/sase-visual/latest-report.json, re-renders that run's manifest with repository and SHA context, uploads the HTML, then re-renders with --report-url "$VISUAL_REPORT_URL" so the summary and annotations point at the freshly uploaded artifact. On execution failure it still uploads available logs and partial captures and reports that failure rather than telling operators to accept goldens. The HTML is uploaded via actions/upload-artifact@v7 with archive: false, which is what makes the per-failure anchors browsable directly from the Actions UI. Expected links point at the immutable https://github.com/<repo>/blob/<sha>/<expected_repo_path> URL; actual/diff links point at the report artifact rather than a public PNG URL because the raw PNGs are only uploaded as a zipped ace-visual-artifacts bundle and have no stable per-file URL.

Add a visual test when the risk is layout, styling, focus highlighting, modal composition, or a regression that is hard to express as state. Prefer a plain state/widget test when the behavior can be asserted through model state, rendered text, selection identity, key handling, or a small widget contract.

Timestamp Display Convention

User-facing timestamp display must go through sase.core.time.parse_local or sase.core.time.format_local, so stored UTC instants, offset-aware values, naive configured-timezone wall times, and epoch values all render in the configured timezone. Naive-model arithmetic keeps using local_now and to_local; storage and wire contracts keep canonical UTC unless their owning schema says otherwise.

tests/test_timezone_display_consistency.py has the focused tz_divergence fixture coverage and the test_no_system_clock_display_sites AST guard. A new bare datetime.now(), argument-less .astimezone(), or tz-less datetime.fromtimestamp() under src/sase/ should normally be fixed by routing through the time helpers instead of adding another guard allowlist entry.

Required Rust Core

Ported sase.core operations are served by the required Rust extension sase_core_rs, distributed as the sase-core-rs package and built from the sibling ../sase-core repo during source development. Normal installs pull a prebuilt wheel; local source installs can build the extension with just install-venv or just rust-install.

There is no pure-Python fallback for ported operations. Use the health check after install changes:

sase core health

See the Rust backend reference for the Python/Rust boundary, shipped Rust-backed operations, source build path, and benchmark expectations.

Linked repositories

sase/sase.yml records linked repositories under repos.linked and sidecars under repos.sidecar. A linked path is relative to the primary checkout, so ../sase-core is the sibling directory next to this repository. sase repo open prepares one repository in a workspace. sase repo list defaults to the current project and the workspace inferred from the current directory. It prints the primary repository, sidecars, these linked repositories, and external repositories already cloned into that workspace. sase repo list --all adds every enabled and disabled project, at the primary workspace. Only sase-core is auto-cloned for every agent launch. The other five stay lazy until sase repo open. plugins.required installs sase-github and sase-research-artifacts from a sibling checkout when one is present, and otherwise from the published package.

Name Path What it is
sase-core ../sase-core Shared Rust core. Auto-cloned for every agent launch. The CI pin is sase-core-revision.txt.
sase-github ../sase-github GitHub VCS and workspace provider for repository, issue, and pull-request workflows. Also a required plugin.
sase-telegram ../sase-telegram Chat-driven workflows and notifications.
sase-nvim ../sase-nvim Syntax, completion, and editor support.
sase-research-artifacts ../sase-research-artifacts @research document provider, research-highlights file hook, and #research* macros. Required plugin; the published package is used when no checkout is present.
sase-listen ../sase-listen Standalone text-to-speech CLI. It turns Markdown into chaptered, loudness-normalized MP3 editions and private podcast feeds.

Source Map

The repository is organized around the CLI entry point, operational subsystems, provider boundaries, and docs/tests:

Path Purpose
src/sase/main/ CLI parser registration and subcommand handlers.
src/sase/ace/ sase's TUI, Patch rendering, query integration, actions, widgets, and TUI state.
src/sase/agent/ Agent launch, detached spawn, prompt fan-out, running-agent metadata, artifact lookup, and naming.
src/sase/axe/ Axe orchestrator, routines, job execution, scheduled jobs, maintenance mode, and automation state.
src/sase/jobs/ Public sase.jobs SDK for axe job scripts (src/sase/chops/ keeps the legacy chop-named facade).
src/sase/macro/ Macro expansion, directives, workflow loading, execution, tracing, explaining, and graphing.
src/sase/macros/ Bundled macro templates, workflows, and schemas shipped with the package.
src/sase/macros/skills/ Bundled agent skill sources and the generated SKILL.md frame.
src/sase/skills/ sase skill CLI helpers, inventory, and use-log implementation.
src/sase/workflows/ Change lifecycle workflows for commit, mentor, CRS, accept, and rewind operations.
src/sase/memory/ Memory inventory, audited read logs, selectors, links, mutation validation, and memory-web operations.
src/sase/core/ Python facade and stable wire records for operations served by sase_core_rs.
src/sase/bead/ Python host layer for bead storage discovery, CLI integration, and epic launch flow.
src/sase/sdd/ Spec-driven development file and bead integration helpers.
src/sase/llm_provider/ Built-in LLM providers and provider registry.
src/sase/vcs_provider/ VCS provider hook specs, plugin registry, and built-in git provider.
src/sase/workspace_provider/ Workspace provider hook specs, plugin registry, and bare-git workspace support.
src/sase/running_field/ Workspace claim and slot-management helpers.
src/sase/procs/ Durable proc store, ids, logs, supervisor, and runner for background work.
src/sase/monitor/ Monitor turn lifecycle: start handoff, detached supervisor, store queries, and follow-up launch.
src/sase/notification_gates/ Command-backed gate bundles, decision receipts, branch execution, and gate CLI helpers.
src/sase/sudo/ Typed sudo gate requests, reviewed terminal handoff, leases, and receipts.
src/sase/notifications/ Notification delivery and storage integration.
src/sase/telemetry/ Local debugging metric accumulation, store queries, health checks, and shared numeric render helpers.
src/sase/version/ Runtime inventory collection and rendering for the sase version CLI command.
src/sase/integrations/ Public helper APIs consumed by external plugins and editors.
src/sase/scripts/ Packaged utility scripts used by axe jobs and support commands.
tests/ Python test suite, with subdirectories mirroring major src/sase/ areas.
docs/ MkDocs Material site source.
sase/sase.yml Repository-local SASE configuration.
sase/task_types.json Committed task-type catalog snapshot written by sase memory init.
sase/macros/ Repository-local macros and workflows for SASE maintenance agents.
sase/memory/ SASE memory files used by repository agents, including generated task_types.md.
sase/repos/ Runtime-only linked, sidecar, and external repository checkouts.
tools/ Development scripts used by just targets and CI checks.

Detailed subsystem pages often include narrower source-layout tables. Use this page for initial orientation, then jump to the specific reference for the area you are changing.

Repository Macros

The checkout's sase/macros/ directory is project-local to the sase repository. When SASE resolves prompts from this project checkout, those entries are namespaced as sase/<name> so they do not collide with user or packaged prompts. Use the catalog's insertion value to know whether an entry should be invoked with # or #!.

Useful visible entries include:

Reference Purpose
#!sase/reads Fan out a reading-recommendation request across Antigravity, Claude, and Codex, then consolidate the final list.
#sase/sync Sync the primary SASE workspace and restart the scheduler.

#!sase/reads accepts a required topic and an optional reference_query. By default, the workflow passes this Dataview query to the research agents:

LIST WITHOUT ID title + " (" + url + ")"
FROM "ref"
WHERE
  source_path AND url AND (
    parent = [[ai_ref]]
    OR parent.parent = [[ai_ref]]
    OR parent.parent.parent = [[ai_ref]]
    OR parent.parent.parent.parent = [[ai_ref]]
    OR parent.parent.parent.parent.parent = [[ai_ref]]
  )
SORT title

Each research agent is expected to use /bob_query to run that query against Bryan's Bob vault, treat every returned title and URL entry as already-known, and only then search for new reading candidates. The three researchers and the final consolidator share one invocation-specific reads-<token> clan, so sase's TUI groups them together while simultaneous invocations remain distinct. A normal invocation can rely on the default query:

#!sase/reads(agent memory systems)

Some repository workflows are marked hidden: true because they are automation helpers, such as docs refresh, recent bug/improvement audits, and Python line-limit splitting. That flag hides workflow run rows in sase's TUI; it does not mean the workflow is unavailable. Use sase macro list or sase's TUI macro browser from a source checkout when you need the exact current catalog.

Documentation Workflow

The docs site is a MkDocs Material project:

Path Purpose
mkdocs.yml Main docs site configuration, strict build, navigation, blog, RSS, and theme settings.
mkdocs-pdf.yml PDF handbook build configuration, inheriting the main site config.
docs/ Markdown, images, stylesheets, JavaScript, redirects, headers, and PDF templates.
site/ Generated site output. It is rebuilt by docs commands and deployed as the static asset directory.

Run the strict site build after changing docs navigation, links, images, or Markdown pages:

just docs-check

Run SASE validation when a change can affect generated initialization files or SDD artifact links. It is deliberately separate from source linting because it can report user/home initialization drift and independently managed SDD state:

sase validate

Run the handbook build and validation when a change materially affects the public handbook, PDF styling, navigation, or generated-site assets:

just docs-pdf-check

just docs-check installs only MkDocs tooling, then runs mkdocs build --strict. just docs-pdf-check installs the PDF tooling, installs Chromium for Playwright, builds mkdocs-pdf.yml in an isolated temporary site directory, post-processes and validates the handbook there, and copies only downloads/sase-handbook.pdf back into site/.

Docs Deployment

Production docs are deployed by .github/workflows/docs-deploy.yml, not by a Cloudflare dashboard build command. The workflow:

  1. Checks out the repo and installs uv, just, and Python 3.12.
  2. Runs just docs-check.
  3. Runs just docs-pdf-check.
  4. Verifies site/index.html, site/_headers, the blog and series pages, and site/downloads/sase-handbook.pdf.
  5. Deploys the prebuilt site/ directory through wrangler.jsonc.
  6. Smoke-tests the deployed handbook PDF from the deployment URL and https://sase.sh/.

The GitHub repository must provide a CLOUDFLARE_API_TOKEN Actions secret with permission to deploy the sase Cloudflare Worker. Keep dashboard-managed Git builds disabled or unused for production so they cannot race the checked in workflow's prebuilt artifact deploy.