Skip to content

sase tui Performance Runbook

This runbook explains how to capture and compare performance data for sase's TUI, the sase tui terminal user interface. It started as the Phase 1 deliverable for the TUI performance overhaul (bead sase-w.1, sdd/epics/202604/tui_perf_overhaul_1.md), and later performance phases still rely on the tracing and benchmark harness described here.

Athena host-relief baseline

Bead sase-zn.1 captured this host-pressure baseline on athena on 2026-09-11 while investigating prompt-input lag in a long-lived sase tui session. Use it as the reference point for the later sase-zn phases that measure remaining CPU, heap, refresh, and scratch pressure.

signal before after host relief
/tmp tmpfs 20G used, 13G available, 62% full 5.6G used, 26G available, 18% full
/ filesystem 837G used, 29G available, 97% full 815G used, 51G available, 95% full
swap 30.4G used of 64G 13.1G used of 64G
sase's TUI process PID 2019865, 10036968 kB RSS, 3332948 kB swap PID 2351038, 766372 kB RSS, 0 kB swap
SASE scratch matches 24 matched /tmp/*cargo-target*, /tmp/*core-target*, /tmp/sase-*-recovery*, and synced managed-root build targets; 37G total 0 remaining matches
SASE_TMPDIR /home/bryan/tmp/sase, via /home/bryan/tmp -> Sync/home/tmp /home/bryan/.cache/sase/tmp in chezmoi-managed .profile and the live sase's TUI process

The cleanup removed only SASE-named cargo/core target directories, SASE recovery bundles, and the stale ~/.sase/perf/tui_trace.jsonl file. The durable environment change is in the chezmoi source home/dot_profile, applied to ~/.profile, then sase's TUI was restarted in tmux pane sase:5.1 after sourcing the updated profile.

Py-spy profiles from the same run are stored outside the repo at:

  • ~/.sase/perf/athena-host-relief-before-20260911T1542.svg
  • ~/.sase/perf/athena-host-relief-after-20260911T1546.svg

Both profile runs reported delayed sampling under remaining host load, so treat the flamegraphs as attribution hints rather than precise wall-clock percentages. The resource deltas above are the reliable phase-1 baseline.

Athena phase-8 verification

Bead sase-zn.8 remeasured the live athena session on 2026-09-12 at 13:15-13:20 EDT, while the host was under real agent load: load average 38.93 / 37.40 / 37.07, many run_agent_runner.py processes were active, and pytest workers from several SASE workspaces were consuming CPU.

signal 2026-09-12 verification
/tmp tmpfs 22M used of 32G; no /tmp/*cargo-target*, /tmp/*core-target*, or /tmp/sase-*-recovery* matches remained
/ filesystem 847G used, 19G available, 98% full
swap 8.3G used of 64G
sase's TUI process PID 54331, up 5h09m; VmRSS: 1199148 kB, VmSwap: 275464 kB, Anonymous: 1163148 kB
sase's TUI CPU main thread sampled at 569 ticks over 60s, about 0.095 cores, or roughly 2.3 CPU-hours/day; whole process was about 0.30 cores under load
SASE scratch live SASE_TMPDIR=/home/bryan/.cache/sase/tmp; managed cargo-targets/ was present and bounded at 68K
old hot frames 25s py-spy raw capture had no reconcile_agent_artifact_index_dismissed_family_members samples and only 4/2399 read_current_notifications_snapshot samples
watchdog not green: the last 30 minutes contained 17 tui_hitch and 5 tui_pump_hitch starts, with fresh Agents-tab hitches during the run
key-to-paint not re-confirmed green: the live sase's TUI process was not started with SASE_TUI_PERF=1; the existing tui_jk.jsonl was stale from 2026-09-10 and showed Agents p95 174.06 ms

The profile is stored outside the repo at ~/.sase/perf/athena-verify-sase-zn8-20260912T131736.raw. py-spy reported sampling delays and 550 read errors under host load, so use the capture for top-frame attribution only. The phase-8 result is therefore mixed: memory pressure, scratch placement, main-thread CPU, and the original reconcile/notification hot frames look bounded, but the watchdog and key-to-paint evidence do not prove the responsiveness target is green. The remaining hitches point at Agents-tab refresh/render work rather than the original sase-zn hot sites.

Agents query load tiering

Bead sase-zu.7 removed the sase-zu load-tiering beta flags on 2026-09-13 and remeasured the not machine:apollo startup cliff. The key behavior is now unconditional: safe indexed query pushdown stays on Tier 1, unsupported queries defer full-history reconciliation instead of blocking first paint, and full-history artifact index loads use the index before falling back to a source scan.

Reproduce the synthetic archive benchmark with the workspace virtualenv:

.venv/bin/python tests/perf/bench_agent_load_tiering.py \
  --output ~/.sase/perf/agent_load_tiering_sase-zu.7_20260913.json

The 2026-09-13 run used sase-core-rs 0.34.24 and a 13,000-artifact fixture with not machine:apollo, 5 measured runs, 1 warmup run, and requested_limit=100.

path / signal before evidence after evidence
startup first paint 2026-09-12 04:57 EDT Tier 1 indexed load: 688 index rows, agents_ready_seconds=8.036 Real archive direct loader from this checkout: Tier 1 artifact-index load, 820 records, 100 returned rows, no error
source-scan cliff 2026-09-12 05:42, 05:56, 06:20 EDT Tier 2 source scans: agents_ready_seconds=23.336, 24.395, 42.663 Synthetic source scan remains the expensive baseline: p50 8621.35 ms, p95 8891.94 ms
bounded indexed first window Regressed path escalated to full source scan for not machine:apollo Synthetic Tier 1 indexed bounded path: p50 107.17 ms, p95 113.48 ms, 113 visible rows
full-history indexed parity Source scan was the only complete-history path for this query Synthetic full-history index path: p50 7672.78 ms, p95 7986.75 ms, missing_count=0, visible_extra_count=0
live log tail 2026-09-13 tui_agent_loads.jsonl still contains old-process full reloads before this checkout is deployed 2026-09-13 tui_startup.jsonl tail shows Tier 1 artifact-index startups, latest agents_ready_seconds=8.013

Interpretation: the measured first-window load is again in line with the indexed 04:57 baseline rather than the 06:20 source-scan cliff. The bounded window is intentionally incomplete (has_more=true) while the complete indexed path proves parity with the source scan. If a future query cannot be narrowed by the artifact index, the loader should report the visible result as incomplete and defer history; it should not increase the startup read to the full archive.

Measured acceptance (sase-zu.8.5)

Bead sase-zu.8.5 remeasured the repaired tree on 2026-09-13 15:00-15:30 EDT: main tree d698f92e05, sase-core-revision.txt 23f19f0 built by just install-venv from the linked checkout (sase-core-rs 0.34.24 dev build, index schema 30, scan wire 9). The sase-zu.7 full-history rows above measured a cached index read, not production's revalidating path, and are superseded by these numbers. The harness now routes every path, the source scan reference included, through the tab's compute_apply_loaded_agents dismissal step, reports read/repair/decode counters, a periodic Tier 1 revalidate, and a settled unchanged-query session:

.venv/bin/python tests/perf/bench_agent_load_tiering.py \
  --fixture-root ~/.cache/sase/tmp/agent-load-tiering-13k --session-refreshes 10 \
  --output ~/.sase/perf/agent_load_tiering_sase-zu.8.5_synthetic13k_20260913.json
.venv/bin/python tests/perf/bench_agent_load_tiering.py --sase-home ~/.sase \
  --output ~/.sase/perf/agent_load_tiering_sase-zu.8.5_athena_real_20260913.json

Both runs used not machine:apollo, 5 measured runs, 1 warmup and requested_limit=100 on a host running other agents. Speedups divide the source-scan p50 by the path's p50; the session baseline models the pre-epic behavior of one source scan per broad refresh.

13,000-artifact synthetic fixture (warm page cache, 270 hidden rows):

path p50 / p95 / max ms read / repair / decode work speedup missing / extra
source scan 8933.0 / 9711.8 / 9711.8 13,000 dirs, 26,003 marker files parsed 1.00x reference
bounded first paint 101.1 / 110.8 / 110.8 0 marker checks, 104 records decoded 88.32x bounded prefix
actual full history 9280.9 / 9928.0 / 9928.0 12,938 marker checks, 0 repaired, 12,803 decoded 0.96x 0 / 0
periodic Tier 1 revalidate 136.0 / 174.9 / 174.9 339 marker checks, 0 repaired, 204 decoded - -
10-refresh settled session 10130.5 total 1 full-history read, 12,938 checks, 13,947 decoded 10.6x -

Athena real archive (11,183 source artifacts):

path p50 / p95 / max ms read / repair / decode work speedup missing / extra
source scan 10866.4 / 11253.6 / 11253.6 11,183 dirs, 41,272 marker files parsed 1.00x reference
bounded first paint 1002.4 / 1202.8 / 1202.8 0 marker checks, 1,062 records decoded 10.84x bounded prefix
actual full history 2988.8 / 4381.5 / 4381.5 7,201 marker checks, 0 repaired, 1,389 decoded 3.64x 13 / 1
periodic Tier 1 revalidate 1087.6 / 1300.6 / 1300.6 6,813 marker checks, 0 repaired, 1,273 decoded - -
10-refresh settled session 10315.0 total 1 full-history read, 7,201 checks, 13,071 decoded 12.6x -

Stage attribution. On the warm synthetic fixture the raw Rust calls are close: source scan p50 1916 ms, cached full-history index read 1624 ms, revalidating full-history read 2283 ms. Revalidation's source-directory reconcile plus 12,938 marker signatures is the epic-added cost (about 660 ms). The remaining time is per-row Python work that both paths share: dict-to-wire conversion (about 1.3 s), Agent decode (about 2.0 s) and the agents-live filter building one field entry per row (about 2.9 s). A single non-selective full-history load is therefore not faster than the source scan on a warm fixture, and this benchmark does not claim it is. On the real archive the raw Rust source scan is 5924 ms against a 1029 ms revalidating index read, because the index drops hidden, dismissed and non-matching rows before decode (956 of 11,183 records). Most of the Tier 1 revalidate cost there is hidden-row repair: about 6,800 marker checks.

Real-archive parity. All 13 missing rows are --code and --mon members of dismissed sessions. The Rust index hides descendants of a dismissed session root (sase-core 34b3229, before this epic), but the tab's Python dismissal step, which the source scan reference uses, does not. The one extra row was an agent created during the run. Before the dismissal step was added, 107 rows differed; every one was a dismissed identity or one of those session members.

Live session: sase's TUI from this tree (sase tui -x --tab agents, PID 2078791, logs in ~/.sase/perf/sase-zu.8.5-tui_{startup,agent_loads,trace}.jsonl) against the real archive, with the not machine:apollo query committed interactively:

signal observation
startup Tier 1 artifact-index startup, on_mount_to_first_paint_seconds=0.414, agents_ready_seconds=6.457
query commit bounded filter load, 1506 ms, 171 rows, has_more=true
automatic history upgrade exactly one input_quiet_tier2_reconcile, 7154 ms, complete_history=true
ordinary refreshes, 14 min 13 cached index auto_refresh loads (p50 1421 ms, max 3181 ms), no repeated full-history load
periodic revalidate 3 tier1_index_revalidate loads, p50 1570 ms
exact deltas 20 watcher, starting-poll and notification deltas, p50 about 150 ms
explicit Full history ,y, then f in the Refresh panel: one manual_full_history load, 4066 ms, complete_history=true
remaining expensive stage 3 bounded loads before the upgrade fell back to a 6.7-7.4 s source scan: artifact index operation lock busy

First paint stays in the indexed baseline range (8.036 s on 2026-09-12; 5.5-6.8 s in the live sase's TUI today). Unchanged covered history is loaded once per committed query. The lock-busy fallbacks come from the process-local index lock's 50 ms read timeout, which 62ad9b657c added. Old artifact directories are outside the bounded startup inotify watch budget. Their marker changes therefore reach the tab through revalidation, not watcher deltas: a touched July done.json showed up in the index as a new done_sig without a rebuild.

Idle-host CPU diet

An idle sase host (TUI open, routines running, no agent work) used to burn roughly four cores: job subprocesses at ~109 spawns/min, each paying ~0.6s of import tax, plus the TUI's unconditional full reconcile every refresh_interval. The idle-CPU-diet epic (sase-wn) turns ticks cheap instead of rare. Use this recipe when a host looks busy at rest again — it should take about five minutes.

Success criteria (idle, no agent work):

signal budget
sase sustained CPU < 0.4 cores
job subprocess spawns < 25/min across every routine
representative job import < 0.2s / < 400 modules
idle TUI < 10% of a core, ~0 stall-watchdog records / 30 min
idle axe collector file opens near zero when nothing changed

The spawn budget sits on a floor no fs trigger can remove: stale_running_cleanup's real input is process liveness (a dead PID touches no file), so it stays on the always trigger in both the hooks (5s) and checks (300s) lanes — ~12/min at full idle on its own. The rest of the shipped-config floor is max_quiet: 120s re-fires from the nine fs-guarded jobs (~4.5/min), the run_every: 30s waits-lane pair (~4/min), and the ≥60s always-trigger lanes (~2/min): ≈ 22/min total. At the post-diet import cost (< 0.2s per boot) that floor is well under a tenth of a core, which is why the budget is a low spawn rate, not zero; a periodic proc-liveness cache could shave the stale_running_cleanup term but has not been worth new machinery.

Deterministic regression floors live in tests, not wall-clock CI:

  • job-SDK import closure: tests/test_chop_import_budget.py (and tests/test_idle_cpu_diet_guardrails.py)
  • zero-spawn idle tick for fs-guarded hooks-lane jobs: tests/test_axe_default_chop_triggers.py / tests/test_idle_cpu_diet_guardrails.py

Five-minute diagnosis

# 1. Fleet spawn rate and no-op ratio (reads routine metrics.json)
sase axe routine status

# 2. Confirm the human overlay matches on-disk counters
jq '{chops_spawned, chops_no_op, chops_skipped, last_tick_spawns, last_tick_skipped,
     spawn_rate_per_minute, no_op_ratio}' \
  ~/.sase/axe/lumberjacks/*/metrics.json

# 3. Process CPU over a quiet minute (/proc deltas)
pidof -x sase; ps -o pid,pcpu,pmem,comm -p $(pgrep -d, -f 'sase (tui|axe)')
# sample twice, 60s apart:
awk '{print $1,$14,$15}' /proc/$(pgrep -n -f 'sase tui')/stat; sleep 60; \
awk '{print $1,$14,$15}' /proc/$(pgrep -n -f 'sase tui')/stat
# utime+stime ticks / (HZ * elapsed) is cores used. On Linux HZ is usually 100.

# 4. Stall-watchdog count over a 30-minute idle TUI session
jq -s 'map(select(.event=="tui_stall" or .event=="tui_hitch")) | length' \
  ~/.sase/logs/tui_stalls.jsonl

# 5. Per-tick TUI counters (surfaces reloaded, axe collector file opens)
SASE_TUI_TRACE=1 sase tui
# quit after a few idle refresh intervals, then:
jq -c 'select(.span=="refresh.auto_tick")' ~/.sase/perf/tui_trace.jsonl | tail
jq -c 'select(.event=="axe.collect")' ~/.sase/perf/tui_trace.jsonl | tail

# 6. Job import cost (wall-clock; the test suite locks the module set, not time)
python -X importtime -c "import sase.jobs.sdk" 2>&1 | tail -1
python -c "import sase.jobs.sdk, sys; print(len(sys.modules))"

# 7. Routine run-history gap analysis (idle ticks should be skipped, not no_op)
#    History lives under the legacy lumberjacks/<routine>/chops/<job>/ paths.
python - <<'PY'
from pathlib import Path
import json
root = Path.home() / ".sase" / "axe" / "lumberjacks"
for routine in sorted(root.iterdir()):
    jobs_dir = routine / "chops"
    for job in sorted(jobs_dir.glob("*")) if jobs_dir.exists() else []:
        index = job / "index.json"
        if not index.exists():
            continue
        ids = json.loads(index.read_text())[:8]
        print(f"== {routine.name}/{job.name} ==")
        for run_id in ids:
            meta = json.loads((job / "runs" / f"{run_id}.json").read_text())
            print(f"  {meta.get('started_at')} {meta.get('status')} {meta.get('reason')}")
PY

A quiet host after the diet should show sase axe routine status Job load at or below the ~22/min shipped-config floor — fs-guarded jobs contribute 0.0 between max_quiet re-fires and their lanes report last tick 0 spawned, while ~12/min of the floor is stale_running_cleanup and is expected — and refresh.auto_tick records with surfaces_reloaded=0 (or only axe/notifications when those tokens actually moved) and axe_file_opens near zero. If spawn rate is well above the floor, check whether shipped jobs lost their fs trigger in src/sase/default_config.yml, or whether ace_refresh_tokens is off.

Suite test-cost gate

The repository-wide pytest cost harness is separate from sase's TUI trace spans. Use it when a change may affect the whole test suite's cost model rather than one TUI interaction:

just test-cost

The lane runs the default fast suite, records per-file wall/CPU time, collection time, the worker RSS curve, and attributed hot causes (now with a per-cause CPU-seconds breakdown alongside wall seconds and count), then prints tools/test_cost_report. Recordings live under ${SASE_HOME:-~/.sase}/test-selection/<project>/timings/cost/; set SASE_TEST_COST_DIR to redirect them. tools/check_test_cost_budgets compares the newest recording with tests/perf/baselines/test_cost_budgets.json, and just check-full plus CI Telemetry's test-cost job enforce that comparison.

Severity model: hard vs. advisory

A single host regularly runs several agents' test suites at once, and Python's time.perf_counter() wall-clock durations stretch under CPU contention that has nothing to do with the suite's own cost -- a causes.ace_page_enter wall time can rise 17% between a quiet host and a busy one for the exact same test run (same invocation count). CPU seconds, invocation counts, and RSS do not move with contention, so the gate is split by metric family instead of failing everything uniformly:

family metrics default tolerance
wall causes.<name> (limit), total_file_wall_seconds, idle_seconds, collection_seconds advisory ci/local
cpu causes.<name>.cpu (cpu_limit), total_file_cpu_seconds, collection_cpu_seconds hard cpu
count causes.<name>.count (count_limit) hard none
memory peak_worker_rss_kib, median_worker_rss_kib, post_collection_worker_rss_kib hard ci/local

Advisory failures are printed with full detail but never fail the gate and are excluded from the exit code -- they are the "the host was probably busy" bucket. Hard failures fail the gate; they are the contention-stable metrics, so an overage there is a real signal. Any budget entry can override its dimension's family default with an explicit enforce (for the wall limit), cpu_enforce (for cpu_limit), or count_enforce (for count_limit) key set to "hard" or "advisory". --suggest always writes these explicitly so the committed file never silently depends on the implicit default.

Count budgets are compared without the cpu/ci/local tolerance: a count_limit already carries ~25% headroom over the worst observed invocation count (round_up_nice(max_observed_count * 1.25)), because counts are near-deterministic -- the headroom belongs in the limit itself, not in a runtime tolerance. That policy trips on a refactor that doubles a call site but tolerates a handful of newly added tests.

CPU attribution note: CostRecorder.measure() times time.process_time(), which is process-wide, so CPU burned by other coroutines on the same event loop during an awaited span is attributed to that span. This matches how wall time already behaves, and within one xdist worker tests run sequentially, so the attribution stays meaningful. Sleep-dominated causes such as pilot_pause_delay report near-zero CPU seconds -- that is correct, and is exactly why the count dimension matters for them.

Run the gate with --strict to make advisories fatal too, for deliberate investigation on a quiet host:

tools/check_test_cost_budgets --strict

just test-cost runs the fast suite and then tools/check_test_cost_budgets (non-strict, so a wall-only advisory no longer breaks the lane). Inside just check-full, tools/run_silent discards a wrapped command's captured output on success, which would otherwise hide a printed advisory; check-full follows the wrapped just test-cost line with an unwrapped tools/check_test_cost_budgets --report-advisories, which re-reads the recording test-cost just wrote and always exits 0, so advisories still reach the operator.

Budget entries are either per-worker or suite-wide totals, and mixing the two up silently defeats the gate:

  • collection_seconds and collection_cpu_seconds are per-worker limits: every xdist worker collects the whole suite, so build_cost_record() sums the metric across workers before a budget entry marked "per_worker": true divides it back down by the record's worker count (falling back to the number of worker payloads that reported collection time, then to 1) before comparing against the limit.
  • Every other summary entry (total_file_wall_seconds, total_file_cpu_seconds, idle_seconds, peak_worker_rss_kib, median_worker_rss_kib, and post_collection_worker_rss_kib) and every causes.* entry is a suite-wide metric, not normalized by worker count.
  • Each worker records start, post_collection, median, and peak RSS summaries. The suite-level worker_rss_curve_kib takes the maximum worker value for start, post_collection, and peak; its median is the median of every positive worker summary value across those four fields, and its sample_count is the sum of worker sample counts. It is an aggregate summary, not one process's time series.
  • peak_worker_rss_kib, median_worker_rss_kib, and post_collection_worker_rss_kib are flat aliases for the corresponding suite-level curve fields. The report and budget suggestion tools do not divide these RSS values by worker count.

Read failures by bucket:

  • total_file_wall_seconds and idle_seconds point to broad suite cost or waiting, but are advisory-only -- pair them with total_file_cpu_seconds (hard) to tell contention from a real regression.
  • collection_seconds, collection_cpu_seconds, and post-collection RSS point to import-time state. Median and peak RSS summarize later retained or run-time growth, but the suite median is over worker summary fields rather than raw time-series samples.
  • Cause entries such as textual_app_run_test_enter, ace_page_enter, parser_create, yaml_load, and subprocess_run point to the hot pattern to audit; each reports up to three failures (causes.<name> wall/advisory, causes.<name>.cpu hard, causes.<name>.count hard).

For focused diagnosis, rerun the cost lane on a path or node and print more rows:

just test-cost -- tests/ace/tui/widgets/test_vim_normal_key_containment.py
tools/test_cost_report --top 20

The committed budgets intentionally have tolerance for host noise. Raise a limit with this workflow, not a hand-picked number, and do not raise one to hide a one-off regression:

just test-cost
tools/check_test_cost_budgets --suggest

--suggest reads the newest retained recordings (--history N to limit the sample), derives each wall/cpu limit as ceil(worst recorded value / (1 + tolerance)) (using the local/ci tolerance for wall metrics and the cpu tolerance for CPU metrics) and each count_limit as ceil(max observed count * 1.25) with no tolerance, rounds each up to a round number, and prints a budget JSON -- with enforce/cpu_enforce/count_enforce written explicitly on every entry -- along with the notes provenance line (sample size, host, UTC date range, worker-count range, node-count range, and per-metric min/median/max) to paste alongside the new limits.

tools/check_test_cost_budgets --ci defaults to os.environ.get("GITHUB_ACTIONS") == "true", not a bare CI variable: an agent runtime that exports CI=true locally must not silently switch a local run onto the wider CI tolerance. GitHub Actions always sets GITHUB_ACTIONS; nothing else does. Pass --ci explicitly to force the CI tolerance on a non-GitHub-Actions host.

Trace recorder

SASE_TUI_TRACE=1 enables tui_trace(...) context managers spread across the Patch, agents, and AXE hot paths. Each entered span emits one JSONL line to:

~/.sase/perf/tui_trace.jsonl

Override the destination with SASE_TUI_TRACE_PATH=/tmp/foo.jsonl. When the env flag is unset the context managers are near-zero-cost no-ops.

Each record contains at least:

ts            unix epoch seconds
span          dotted span name (e.g. "agents.refresh_panel_widgets")
duration_ms   wall time inside the span
current_tab   "artifacts" | "agents" | "services" | null

…plus any per-call counters (count, agents, panels, output_bytes, …) and any global context fields seeded via sase.ace.tui.util.trace.set_trace_context(...) (the app pushes current_tab and current_idx automatically).

Point-in-time records emitted by trace_event(...) contain event instead of span/duration_ms. They are used for selection and highlight watcher transitions where there is no timed block to measure, and for the prompt-submit pipeline: launch.accepted fires in the handler that unmounts the prompt bar, and launch.submitted fires when the durable sase run proc takes over, carrying accept_to_submit_ms and the pending-launch stages it waited in. Pair them by launch_id (jq -c 'select(.event|startswith("launch."))' ~/.sase/perf/tui_trace.jsonl) to measure how long a submit stayed pending after the bar disappeared.

Timed spans for the main Patch, agents, and AXE hot paths (by file, relative to src/sase/ace/tui/; run rg 'tui_trace\(' src/sase for the complete set, which also covers loader stages and pager opens):

  • actions/patch/_display.py — patch.refresh_display, patch.refresh_debounced, patch.refresh_detail_only
  • actions/patch/_loading.py — patch.filter
  • actions/agents/_display.py — agents.refresh_display, agents.refresh_display_incremental, agents.refresh_debounced
  • actions/agents/_display_panel_widgets.py — agents.refresh_panel_widgets
  • actions/agents/_display_panel_layout.py — agents.refresh_panel_highlights, agents.refresh_focused_panel
  • actions/agents/_loading_helpers.py — agents.load_from_disk
  • actions/agents/_loading_live_hints.py — agents.live_hint_refresh
  • actions/agents/_display_detail_render.py — agents.view_hints_refresh
  • actions/hints/_files.py — agents.view_files, agents.view_agent_files, agents.view_hint_bar_mount
  • widgets/prompt_panel/_agent_display_hints.py — widget.prompt_panel.update_display_with_hints
  • widgets/patch_list.py — widget.patch_list.update_list, widget.patch_list.update_highlight, widget.patch_list.patch_patch_row
  • widgets/patch_detail.py — widget.patch_detail.update_display
  • widgets/agent_list.py — widget.agent_list.update_list, widget.agent_list.update_highlight, widget.agent_list.patch_agent_row, widget.agent_list.try_remove_rows
  • widgets/agent_detail.py — widget.agent_detail.update_display, widget.agent_detail.update_display_immediate
  • widgets/artifacts/relation_panel.py — widget.relation_panel.update_relations
  • widgets/prompt_panel/_agent_display.py — widget.prompt_panel.update_display, widget.prompt_panel.update_header_only, widget.prompt_panel.update_tribe_display, widget.prompt_panel.refresh_slow_tool_metadata_from_cache
  • widgets/prompt_panel/_agent_display_header_summary.py — widget.prompt_panel.build_detail_header_summary and one child span per resolver (see "SASE CONTEXT enrichment" below)
  • widgets/file_panel/_panel.py — widget.file_panel.update_display
  • widgets/llm_calls_panel.py — widget.llm_calls_panel.update_display
  • widgets/axe_dashboard.py — widget.axe_dashboard.update_display, widget.axe_dashboard.update_lumberjack_overview (routine overview), widget.axe_dashboard.update_chop_run_display (job detail)

Spans nest cleanly: a single keypress that fires agents.refresh_debounced will record one outer span plus inner widget.agent_list.update_highlight and agents.refresh_panel_highlights spans.

Three agents-tab spans carry counters meant for soak assertions on tribe-panel stability:

  • agents.finalize_query_filter reports agents_in (the filter's input roster) and agents_out (its output). The old single agents counter was the input, which is easy to misread as the published roster size.
  • agents.apply_loaded_agents_prepared reports finalize_plan as applied, discarded, or absent. A discard also carries finalize_plan_discard_reason (stale_token or roster_fingerprint). A soak that never sees applied proves nothing about the off-thread finalize path: every apply fell back to the inline one.
  • agents.refresh_panel_widgets reports panel_widget_ids, the tribe-stable widget ids in mount order, so a tribe panel's continuous presence is assertable from the trace.

Agents-tab paint frames (agents.paint_frame)

Counters such as the ones above pass on an idle tab, so they cannot show what a panel looks like while a node joins it. With SASE_TUI_TRACE=1 every completed Agents display refresh (and every out-of-band change of the agent-list column width) also emits one agents.paint_frame event: kind (full_rebuild, incremental, highlight, container_width), the refresh's source / display_cost / fallback_reason, the app's app_grouping_mode, the selected_identity, #agent-list-container's negotiated container_width (its styles.width) beside the laid-out container_layout_width, and a panels list. Each panel entry carries widget_id, object_id (the widget's Python id(), so a remount is visible), option_count, collapsed, collapse_intent, grouping_mode, requested_width, height, highlighted, highlighted_identity, scroll_y and viewport_height. seq orders the frames.

# One row per frame: what changed on screen between consecutive refreshes.
jq -c 'select(.event == "agents.paint_frame")
       | {seq, kind, source, display_cost, fallback_reason, container_width,
          panels: [.panels[] | [.widget_id, .option_count, .requested_width, .height]]}' \
   ~/.sase/perf/tui_trace.jsonl

A node joining @epic should produce exactly one frame whose panels or container_width differ from the previous one. A full_rebuild frame carrying fallback_reason: stale_grouping_mode, or a container_width frame right after a refresh frame that already widened a panel's requested_width, is the flicker itself. tests/ace/tui/test_epic_panel_arrival_frames.py asserts these invariants deterministically; the paint log lives in actions/agents/_paint_log.py.

The arrival's display_cost says which path painted it. An ordinary node (no clan, session or workflow relation, no wider than the panel's existing rows, no new banner) records display_row_insert: the row is inserted in place and no update_list runs. A node that cannot be inserted still rebuilds only its own panel, recording display_panel_rebuild after a display_row_insert record whose fallback_reason names the gate that declined it (width_growth, status_membership_change, workflow_tree_change, panel_membership_change). A STARTING agent is not rendered, so the apply that first shows it is the arrival, not a status-bucket move. Tally the paths over a soak:

jq -r 'select(.event == "agents.paint_frame" and .kind != "settled")
       | [.display_cost, .fallback_reason] | @tsv' ~/.sase/perf/tui_trace.jsonl \
   | sort | uniq -c | sort -rn

A soak that never creates a node shows no display_row_insert at all; that proves nothing about arrivals.

Rebuild scope is decided per panel. When a roster change concerns only some panels (a duplicated identity, a BY_STATUS bucket change, or a change to a panel's workflow tree), the apply rebuilds those panels and patches the rest: display_panel_rebuild on an incremental frame, never display_full_rebuild. Each panel it names is one agents.refresh_work event with stage: display_fallback, the fallback_reason (panel_membership_change, status_membership_change, workflow_tree_change or group_fold_change), and panel, that panel's widget id. A group_fold_change rebuild means only a grouping-banner fold changed: the roster is identical, but the panel's rows were built under different fold state. Tally which panels a soak rebuilt, and why:

jq -r 'select(.event == "agents.refresh_work" and .stage == "display_fallback"
              and .panel != null)
       | [.panel, .fallback_reason] | @tsv' ~/.sase/perf/tui_trace.jsonl \
   | sort | uniq -c | sort -rn

A display_full_rebuild therefore means the whole tab was rebuilt: a changed search query, an unsupported or stale grouping mode, or a fast path that refused the apply.

To look at the same arrival by eye, capture the @epic panel before and after a node joins it with a live sase screenshot. This is a human check and stays out of the golden lane (live captures carry real timestamps and host state):

# Before: capture a fresh TUI on the Agents tab and keep its tmux window.
sase screenshot --keep -o /tmp/epic_before.png -w "@epic" -- -t agents
# Copy the printed sase_tmux_target=..., launch a cheap agent that lands in the epic
# tribe (a prompt tagged `#tribe:epic`), wait for its row to appear, then:
sase screenshot --window <sase_tmux_target> -o /tmp/epic_after.png

The tmux-launched TUI that sase screenshot starts gets SASE_TUI_TRACE=1 injected automatically, so its frame events land in ~/.sase/perf/tui_trace.jsonl to compare.

Compare the two PNGs for the panel border, scroll position, highlighted row and column width, and confirm the same arrival in the agents.paint_frame events above. Do not copy either PNG under tests/ace/tui/visual/snapshots/png/; goldens are maintained only by just fix-tui-screenshots.

sase's TUI deliberately keeps live-workspace pencil hints off the startup-critical agents loader. The first load classifies only cheap persisted diff_path badges. After that agents list has applied, agents.live_hint_refresh runs VCS probes for active, non-terminal rows that do not yet have a persisted diff and patches changed rows in place. During startup investigations, treat agents.load_from_disk and agents.live_hint_refresh as separate costs: the former controls time to first interactive Agents-tab paint, while the latter explains deferred pencil-badge updates.

Reading a view-hints capture

Pressing v on the Agents tab nests four spans, outermost first:

agents.view_files                              whole v keypath (both tabs)
└─ agents.view_agent_files                     Agents-tab branch only
   ├─ widget.prompt_panel.update_display_with_hints    the annotated render
   └─ agents.view_hint_bar_mount               mounting the HintInputBar

agents.view_files is the keypress → hint-bar-mounted interval: today the bar is mounted last, so this span is effectively the render cost plus the mount cost. Subtracting agents.view_hint_bar_mount from it gives the part of the wait the user pays for work that is not the bar appearing.

agents.view_hints_refresh is the same annotated render fired again from an Agents-tab detail repaint or from the detail-header enrichment message, rather than from a keypress. Seeing it repeatedly with hint mode active is the signal that the document is being rebuilt on refresh.

Useful counters:

agents.view_files          tab
agents.view_agent_files    agent_session_container, hints, commit_views,
                           header_enrichment_pending, outcome
                           (mounted | refocused | empty | no_agent |
                            detached_container)
agents.view_hints_refresh  agent_session_container, hints, commit_views
update_display_with_hints  agent_session_container, hints, commit_views,
                           tool_call_reports, annotated_chars,
                           header_summary (warm | cold)

annotated_chars counts every character handed to the hint scanner, summed across fragments and session members, so it is the size term to divide a duration by. header_summary says whether the render had a warm detail-header summary: a cold render omits the SASE CONTEXT hints entirely and will be rebuilt when the enrichment worker lands.

Slice one press out of a capture with:

jq -c 'select(.span | startswith("agents.view_") or . == "widget.prompt_panel.update_display_with_hints")
       | {span, duration_ms, hints, annotated_chars, header_summary, agent_session_container, outcome}' \
   ~/.sase/perf/tui_trace.jsonl | tail -20

Trace events currently wired include:

  • selection.current_idx.set
  • widget.patch_list.watch_highlighted and .suppressed
  • widget.agent_list.watch_highlighted and .suppressed
  • widget.bgcmd_list.watch_highlighted and .suppressed

SASE CONTEXT enrichment (detail-header summary)

build_detail_header_summary (widgets/prompt_panel/_agent_display_header_summary.py) is the worker that resolves the PLAN, BEAD, ARTIFACTS, MEMORY, SKILLS, and WORKSPACES lanes of the SASE CONTEXT section (bead sase-l6.1, plans/202608/sase_context_incremental.md). It emits one parent span, widget.prompt_panel.build_detail_header_summary, plus one child span per resolver, all sharing that dotted prefix:

widget.prompt_panel.build_detail_header_summary                 (parent)
  .macros_used
  .bead_display
  .plan_enrichment
  .slow_tool_sources
  .agent_page_url
  .linked_delta_groups
  .artifact_file_paths        the one resolver with no cache — usually the
                               most expensive lane before phase `stores` lands
  .artifact_reads
  .memory_reads
  .skill_uses
  .opened_workspaces
  .delta_entries
  .wait_bead_statuses

The parent span carries agent (the agent's cl_name) and cache_state ("cold" on the first time this process has resolved that agent identity, "warm" after). This marker is process-local and best-effort — it tracks whether this worker has touched the identity before, independent of and coarser than each resolver's own on-disk cache — so it exists to make a raw capture readable without cross-referencing every resolver's cache state.

Since phase stream (bead sase-l6.4), one selection's enrichment worker resolves its requested lanes cheapest-first across LANE_RESOLUTION_BATCHES (_agent_display_header_summary.py) and merges/publishes each batch as it lands, so a single selection now emits up to three widget.prompt_panel.build_detail_header_summary parent spans back to back — one per non-empty batch — instead of one. All three share the same agent; only the first carries cache_state: "cold" for a never-before-seen identity; the batches after it reuse the same process-local seen-set and read "warm". Group by agent and read the spans in capture order to see which lane group landed first (typically the free/cached-lookup batch) versus last (typically the store-backed ARTIFACTS/MEMORY/SKILLS batch, before phase stores's caches are warm).

Slice one selection's enrichment out of a capture with:

jq -c 'select(.span | startswith("widget.prompt_panel.build_detail_header_summary"))
       | {span, duration_ms, agent, cache_state}' \
   ~/.sase/perf/tui_trace.jsonl | tail -20

Reproduce the baseline table (per-resolver cold/warm cost, plus where artifact_file_paths spends its time inside list_artifact_files) with the committed benchmark, which drives the exact same spans read above rather than re-timing the resolvers by hand:

pytest -s -m slow tests/perf/bench_detail_header_summary.py

python -m tests.perf.bench_detail_header_summary --include-home \
    --count 20 --output ~/.sase/perf/detail_header_summary_baseline.json

--include-home is required to measure real ~/.sase agents; without it the script only exercises a tiny hermetic in-memory fixture, so it stays safe to run in CI. --count controls how many non-clan agents from load_tiered_agents(full_history=False) are sampled.

Heap sampler

SASE_TUI_HEAP=1 enables an opt-in tracemalloc sampler for long-lived sase's TUI sessions. Tracing starts when AceApp initializes, while each snapshot is scheduled from the TUI timer into a pump-free task and written from a worker thread. Samples append one compact JSONL record to:

~/.sase/perf/tui_heap.jsonl

Override the destination with SASE_TUI_HEAP_PATH=/tmp/tui_heap.jsonl. The default interval is 300 seconds; override it with SASE_TUI_HEAP_INTERVAL_SECONDS. One sample is also written at startup. Each record includes current_bytes, peak_bytes, and the top allocation sites grouped by source line, each with its captured traceback. The sampler records 25 sites with 25 traceback frames by default; tune them with SASE_TUI_HEAP_TOP_N and SASE_TUI_HEAP_NFRAME. It keeps no previous snapshots in memory, so use adjacent JSONL records to compare growth over time.

Quick capture

SASE_TUI_TRACE=1 sase tui
# … exercise the path you care about (cold start, query change, j/k burst,
#   auto-refresh, large reply select) …
# Quit with q.

# Inspect:
jq -c 'select(.span | startswith("widget.agent_list."))' \
   ~/.sase/perf/tui_trace.jsonl | head -20

To inspect point events instead of timed spans:

jq -c 'select(.event)' ~/.sase/perf/tui_trace.jsonl | head -20

For key-to-paint timing during j/k navigation, also enable the separate perf recorder:

SASE_TUI_TRACE=1 SASE_TUI_PERF=1 sase tui
jq -c . ~/.sase/perf/tui_jk.jsonl | head -20

Override the key-to-paint path with SASE_TUI_PERF_PATH=/tmp/tui_jk.jsonl.

Agents that launch the TUI via sase tui --tmux get SASE_TUI_TRACE=1 and SASE_TUI_PERF=1 injected automatically; export the variable to 0 before invoking to opt out. SASE_TUI_HEAP is never auto-enabled; set it explicitly for heap attribution.

Offline Fleet fault j/k benchmark

Use the committed offline Fleet fault benchmark when changing unified remote projection, federation refresh, remote lifecycle row metadata, or Agents navigation around remote rows:

just test-slow tests/ace/tui/bench_tui_jk.py -k fleet_jk_fault_scenarios

The benchmark uses in-memory federation fixtures only: no enrolled machines, network access, or real federation worker are involved. It warms the rendered remote rows, then records Agents-tab j/k key-to-paint samples while each fault script is active:

  • hung_host keeps a healthy host's rows while a second host contributes a deadline diagnostic.
  • reconnect_churn keeps rows visible while one host reports aging/reconnecting health.
  • event_burst applies a larger multi-host catalog burst.

Read the printed p50/p95/max tables per scenario. Every next and prev p95 must be < 16 ms, and the run must leave no tui_stalls.jsonl rows. A p95 failure points at the Agents remote-row key-to-paint path; any stall row means the event loop or Textual pump was blocked and must be investigated before landing.

Unread and idle baseline (epic sase-1d7)

Use the committed unread benches when changing unread acknowledgment, unread navigation, the 1 Hz runtime tick, or fleet reprojection. They record the baseline the later epic phases compare against:

just test-slow tests/ace/tui/bench_tui_jk.py -k "unread_bulk_ack or unread_jump"
just test-slow tests/perf/bench_tui_trace.py -k idle_tick_and_fleet_refresh

Both benches drive a screenshot-shaped roster (about 200 agents, clan containers, three tribes) with in-memory fixtures only: the notification dismiss write is stubbed, so no real notifications are touched. The unread bench covers ,u and ,j with the target visible, inside a collapsed panel, inside a collapsed clan, and on another Agents tab; the idle bench times one countdown tick (_patch_agent_runtime_rows) and one fleet_refresh reprojection (forced and unchanged variants) by direct call, because the tick and fleet apply are skipped while navigating and the j/k benches cannot see them. The committed ceilings are deliberately generous (order-of-magnitude regression only); compare the printed p50/p95/max tables against the bead-note baseline instead.

The same spans can be captured live from a real session:

SASE_TUI_TRACE=1 SASE_TUI_PERF=1 sase tui
# ... press ,u / ,j /,J on the Agents tab, then quit with q ...
jq -c 'select(.span == "leader.unread_bulk_ack" or .span == "leader.unread_jump"
  or .span == "unread.ack_complete" or .span == "unread.reconcile"
  or .span == "unread.chrome_apply"
  or .span == "agent_nodes.projection_index")' ~/.sase/perf/tui_trace.jsonl
jq -c 'select(.action == ",u" or .action == ",j" or .action == ",J")' \
  ~/.sase/perf/tui_jk.jsonl

Note: an off-tab ,j that switches Agents tabs records its paint under the agents_tab_switch action, not ,j — the switch starts its own sample. The store_bytes counter on unread.ack_complete is statted on the ack worker thread; the UI thread never stats the store itself.

Freeze and hitch capture

sase's TUI always-on watchdog writes event-loop and Textual message-pump diagnostics to ~/.sase/logs/tui_stalls.jsonl. There are two independent severity tiers:

  • tui_hitch / tui_pump_hitch fire after 1.5 seconds by default. These are compact records containing the current tab and selection, the last action, keypress age, and the main-thread stack. Recovery rows use the corresponding *_recovered event and include the episode duration. Hitch records deliberately omit full asyncio-task and worker thread dumps, and are rate-limited; suppressed_count reports episodes omitted since the previous admitted record.
  • tui_stall / tui_pump_stall retain the existing 5-second threshold and richer task/thread diagnostics. A long freeze can produce both a hitch and a stall because the two state machines are independent.

Lower both tiers for a short verification soak with:

SASE_TUI_HITCH_THRESHOLD_SECONDS=0.25 \
SASE_TUI_PUMP_HITCH_THRESHOLD_SECONDS=0.25 \
SASE_TUI_STALL_THRESHOLD_SECONDS=0.75 \
SASE_TUI_PUMP_STALL_THRESHOLD_SECONDS=0.75 \
SASE_TUI_STALL_POLL_INTERVAL=0.02 \
SASE_TUI_PUMP_STALL_POLL_INTERVAL=0.02 \
SASE_TUI_STALL_PATH=/tmp/sase-tui-soak.jsonl \
sase tui

Exercise startup typing, launch and cleanup bursts, prompt-history and revive-agent modal opens, tab switching during agent churn, and refreshes while the artifact or dismissed-bundle indexes are contended. A fixed path should stay interactive and produce no hitch/stall rows. Inspect any records in chronological order:

jq -c '{event, stall_seconds, duration_seconds, current_tab, current_idx,
        last_action, last_keypress_age_s, suppressed_count, main_thread_stack}' \
  /tmp/sase-tui-soak.jsonl

Read main_thread_stack from its final frame upward to find the blocking call. For a pump-only event, the loop may still be running while one Textual message pump is awaiting slow work; a loop event means the event loop itself stopped servicing the watchdog beacon. Disable the compact tiers independently with SASE_TUI_HITCH_DISABLE=1 and SASE_TUI_PUMP_HITCH_DISABLE=1; the existing stall-tier disable flags remain SASE_TUI_STALL_DISABLE=1 and SASE_TUI_PUMP_STALL_DISABLE=1.

Hitch and stall rows (and their recoveries) carry additive whole-process fields:

  • late and detected_by — true / "watchdog_lateness" when the watchdog's own poll was late even though the loop beacon had already run (the GIL-race blind spot: a stop-the-world pause on another thread freezes every thread, and the beacon recovers first). On-time rows report false / "loop_gap" (or "pump_gap").
  • poll_lag_s — the poll lateness minus one poll interval (0.0 on on-time rows).
  • net_stall_seconds — the duration minus one poll interval, floored at the tier threshold. stall_seconds / duration_seconds are unchanged.
  • app_instance_id — the per-instance ID minted at app construction, so a busy hour spanning restarts splits cleanly per instance.
  • Recovery rows add gc_overlap_s (GC pause seconds overlapping the episode), gc_generations, and gc_triggers, attributed from the GC telemetry ring. A hitch caused by an intentional idle collection is then distinguishable from an interactive freeze.

Filter one instance's late freezes with:

jq -c 'select(.event == "tui_hitch" and .late == true) | {ts, stall_seconds,
        poll_lag_s, net_stall_seconds, gc_overlap_s, app_instance_id}' \
  ~/.sase/logs/tui_stalls.jsonl

The persistent TUI diagnostic JSONL files under ~/.sase/logs/—stall, git-operation, launch-timing, external-tool, agent-load, and startup records—rotate independently before appending a record would make a non-empty file exceed 2 MiB. Each keeps one .1 generation. Set SASE_TUI_TELEMETRY_MAX_BYTES to another per-file byte limit, or 0 for no size rotation. This bound is separate from the opt-in trace files under ~/.sase/perf/.

GC pause and memory heartbeat rows

The GC telemetry service (src/sase/ace/tui/util/gc_telemetry.py, installed from _start_post_first_paint_services) adds two more row kinds to tui_stalls.jsonl:

  • tui_gc_pause — one row per gen-2 collection and per collection ≥ 50 ms, with generation, duration_s, thread, trigger ("automatic" unless the collection ran under gc_trigger(...), e.g. an intentional idle collection), collected / uncollectable, and app_instance_id. Rows are rate-capped at 60/min; suppressed_count reports pauses folded into a row since the previous admitted one. Heartbeat totals stay exact either way.
  • tui_memory_heartbeat — every 5 minutes (plus once ~15 s after startup), with rss_bytes / vmswap_bytes (from /proc/self/status, null off Linux), major_faults and major_faults_delta (from /proc/self/stat), gc_count / gc_threshold / gc_freeze_count, exact per-generation count / total_s / max_s for the window they cover (gc_generations), instance uptime_s (since install), and window_s (wall-clock span since the previous heartbeat, so GC share is exact as total_s / window_s).

Every row carries app_instance_id (also on tui_startup and tui_agent_load rows), so a busy hour spanning restarts splits cleanly per instance. Filter one instance's GC share with:

jq -c 'select(.event == "tui_gc_pause" and .app_instance_id == "<id>")
  | {generation, duration_s, trigger, thread}' ~/.sase/logs/tui_stalls.jsonl

Disable with SASE_TUI_GC_TELEMETRY_DISABLE=1. The service never auto-installs under the sase.ace.testing harness.

The watchdog contributes exact hitch totals to each heartbeat through a stall_watchdog heartbeat provider: loop_hitch_episodes / loop_hitch_seconds and pump_hitch_episodes / pump_hitch_seconds count every recovered episode in the window, while loop_suppressed_episodes / loop_suppressed_seconds (and the pump equivalents) count the rate-limited ones. Row counts stay bounded; the totals stay exact.

GC idle-collection policy

The idle-GC policy (src/sase/ace/tui/util/gc_policy.py, installed from _start_post_first_paint_services right after telemetry) keeps automatic gen-2 collections off the interactive path:

  • Once startup loads settle (_mount_state_loads_done) and input first goes idle, one gc.collect() followed by gc.freeze() runs under the startup_freeze trigger, exactly once per app instance. It is never re-frozen: cyclic garbage among frozen objects is never reclaimed.
  • On the classic three-generation collector the policy preserves the current t0/t1 and raises threshold2 to 10_000, so automatic full collections become a rare backstop (about once an hour at 2.9 gen-1/s). Any other collector shape leaves thresholds alone and records threshold_policy: "unsupported" (instead of "raised") on the tui_memory_heartbeat row; the idle collector still runs.
  • A 1 s tick runs gc.collect() on the UI thread only when input has been quiet for 3 s, no prompt is active, the navigation gate is idle, no modal screen is accepting text input, at least 60 s have passed since the last full collection of any trigger, and a collection is due (100 gen-1 collections since the last full one). Mouse clicks and scrolls count as input for the quiet window without consuming the events.
  • After 5 minutes without a full collection, or when RSS grew more than 500 MB since the last full one (from the heartbeat's cached sample, never a UI-thread /proc read), the quiet window relaxes to 1 s. These run tagged backstop / rss_backstop; the raised threshold2 remains the hard backstop.
  • After an idle or backstop collection, freed arenas are released with malloc_trim on a daemon worker thread (sase-tui-gc-trim, via the trim_allocator half of sase.axe.runner_idle_memory), at most every 10 minutes.

Confirm it from tui_gc_pause rows: intentional collections carry trigger values startup_freeze / idle / backstop / rss_backstop, while anything still tagged "automatic" started on its own. An idle collection that overlaps a hitch is attributable through the same trigger. Disable freeze, threshold, and the idle collector with SASE_TUI_GC_POLICY_DISABLE=1 (telemetry stays on). The policy never auto-installs under the sase.ace.testing harness, and teardown restores the original thresholds, calls gc.unfreeze(), and stops the tick timer.

Frozen-share report

tools/tui_freeze_report is the read-only measuring tool for the freeze epic. It reports per app_instance_id (rows that predate instance IDs are bucketed into tui_startup windows instead of being dropped) over a --since / --until window:

tools/tui_freeze_report --since 2026-10-01T12:00 --until 2026-10-01T13:00
tools/tui_freeze_report --instance abc123 --path /tmp/sase-tui-soak.jsonl

Each instance reports the loop-only union frozen share (pump and loop tiers are never double-counted), hitch duration median / p90 / max, the late versus on-time split, GC share by generation and trigger, idle-collection pause p50 / p95 / max, automatic gen-2 collections within 2 s of input (a documented heuristic: it needs a last_keypress_age_s context row within 2 s of the pause), and the RSS / swap trajectory from heartbeats.

Startup telemetry capture

~/.sase/logs/tui_startup.jsonl (sase/logs/tui_telemetry.py:log_tui_startup) gets one durable record per sase's TUI session, written after both the Agents and Services surfaces finish their first load — the same point the visible startup-stopwatch badge stops. It exists so every "startup dropped from X to Y" claim in this repo's plans and epics is checkable against a real terminal run instead of a modelled component sum (plans/202608/ace_startup_critical_path.md).

Each record carries two headline metrics, both measured from the App's on_mount — the same anchor the visible stopwatch badge uses, so the two numbers should visually track what you saw on screen:

  • all_surfaces_ready_seconds — elapsed time until both the Agents and Services tabs' first load finished, regardless of which tab was visible. This is today's stopwatch semantics.
  • visible_ready_seconds — elapsed time until the initially visible tab's own surface was interactive. Recorded from day one even though nothing currently drives the stopwatch off of it, so a later change to end the stopwatch on the visible surface instead of both surfaces has an honest before/after rather than a redefinition that silently invalidates every prior capture.

The record also carries process_start_to_on_mount_seconds and on_mount_to_first_paint_seconds (the two stages before either surface starts loading), agents_ready_seconds / axe_ready_seconds (each surface's own elapsed time from process start), initial_tab, agent_row_count, index_row_count (the Tier-1 index query's row count when the load went through the persistent index, null otherwise), and source / tier / artifact_source from the load's AgentLoadState.

process_start_to_on_mount_seconds keeps its historical meaning (the AceApp init stamp to on_mount). Four additive fields split the wider OS-process-start → compose path without changing that field:

  • interpreter_cli_import_seconds — OS process start to just before AceApp import
  • app_module_import_seconds — from sase.ace.tui import AceApp
  • app_construct_seconds — AceApp() construction
  • compose_seconds — Textual consumption of compose()

Missing stamps record as null (unit tests and non-CLI constructions).

Attributed startup capture (sase-132.1)

One command sequence reproduces a fully attributed capture at a recorded SHA. Store outputs under ~/.sase/perf/ with the epic prefix and label the host as quiet or busy from load average plus run_agent_runner.py count (see the script's host_state).

# 1. Standalone loader bench against the real archive (label host state)
.venv/bin/python tests/perf/capture_tui_startup.py bench \
  --sase-home ~/.sase \
  --output-dir ~/.sase/perf

# 2. Three live traced startups + one pyinstrument profile
.venv/bin/python tests/perf/capture_tui_startup.py live \
  --runs 3 --profile \
  --output-dir ~/.sase/perf

# 3. Import-time of the TUI app module in this environment
.venv/bin/python tests/perf/capture_tui_startup.py importtime \
  --output-dir ~/.sase/perf

Or all three:

.venv/bin/python tests/perf/capture_tui_startup.py all --output-dir ~/.sase/perf

Query startup-window contention without reconstructing the window from timestamps:

jq -c 'select(.startup_window==true)' ~/.sase/perf/sase-132.1-*/tui_trace.jsonl

Loader wait-vs-work inside the startup agents.load_from_disk span:

jq -c 'select(.span=="agents.load_from_disk"
        or .span=="agents.load_from_disk.dismissed_snapshot"
        or .span=="agents.load_from_disk.provider"
        or .span=="agents.load_from_disk.index"
        or .span=="agents.load_from_disk.decode"
        or .span=="agents.load_from_disk.projections")' \
  ~/.sase/perf/sase-132.1-*/tui_trace.jsonl

Axe first load (duration span plus the legacy point event, both with file_opens):

jq -c 'select(.span=="axe.startup" or .span=="axe.load_status"
        or .span=="axe.collect" or .event=="axe.collect")' \
  ~/.sase/perf/sase-132.1-*/tui_trace.jsonl

To capture a before/after pair:

# baseline, before your change
git stash  # or check out the prior commit in a second workspace
sase tui   # quit once the tabs finish loading; repeat 3x for a stable read
git stash pop

# after your change
sase tui   # quit once the tabs finish loading; repeat 3x

Then compare the two sets of records:

jq -c '{timestamp, all_surfaces_ready_seconds, visible_ready_seconds,
        agents_ready_seconds, axe_ready_seconds, agent_row_count,
        index_row_count, source, tier, artifact_source,
        process_start_to_on_mount_seconds,
        interpreter_cli_import_seconds, app_module_import_seconds,
        app_construct_seconds, compose_seconds}' \
  ~/.sase/logs/tui_startup.jsonl

tui_agent_loads.jsonl's slow-stage records are censored below _DEFAULT_SLOW_LOADER_STAGE_THRESHOLD_SECONDS (2.0 s in src/sase/ace/tui/actions/agents/_loading_disk_support.py); a capture run against a tree that has already dropped under 2 s needs the sub-threshold stages too. Override the threshold for the run instead of editing the constant (non-numeric or non-positive values fall back to the default):

SASE_TUI_LOADER_LOG_THRESHOLD_SECONDS=0.05 sase tui

TUI import-budget ratchet (sase-13p)

tests/ace/tui/test_app_import_budget.py guards the fresh-interpreter module count for import sase.ace.tui.app with a strict < against _MAX_MODULE_COUNT (measured plus 20). The cap only moves down: raising it requires a named module that is genuinely needed on the first-paint path with deferral impossible, recorded in the test comment with its commit, and the raise covers only the attributed amount. Never redefine success as <= and never raise the cap to absorb unattributed drift. New TUI code must not add eager imports to the startup closure; use use-site or TYPE_CHECKING imports and extend the test's deferred_modules probe with the deferred heavy modules.

Measure and attribute with the closure tool (run it with the project venv):

.venv/bin/python tools/tui_import_closure
.venv/bin/python tools/tui_import_closure --diff $(git merge-base HEAD origin/master)

The bare form prints the current closure count. --diff REV measures the same import at REV from a temporary detached worktree (removed afterwards) and lists the sase.* modules added and removed.

Synthetic-data benchmark harness

The harness lives at tests/perf/bench_tui_trace.py. It generates in-memory Patch and agent fixtures, then drives the TUI through Textual Pilot without touching real ~/.sase data. It is marked pytest.mark.slow, so it does not run as part of just test.

For the Updates > Plugins catalog-scale path, use the focused benches and the committed measuring-stick baseline:

just bench-plugin-catalog-scale

That records p50/p95/max at 10 / 250 / 1000 / 2000 entries for pane open, one filter keystroke, 20 j presses (through the queued OptionHighlighted handler), ' jump-hint allocation, and one I mark toggle, plus non-TUI enrich/fetch cost curves. The committed numbers live in tests/perf/baselines/plugin_catalog_scale_baseline.json. Filter-keystroke and j-press p95 must stay under 16 ms at n=2000; eager enrich fetches must stay O(installed) and enrich's scan_work (catalog rows walked, counted through CountingEntries) must stay linear in n. Rewrite the recorded rows after a capture run with python -m tests.perf.bench_plugin_catalog_scale --write-baseline and SASE_PLUGIN_CATALOG_SCALE_WRITE_BASELINE=1 pytest -s -m slow tests/ace/tui/bench_plugins_catalog_scale.py. just plugin-catalog-scale-check is the CI regression floor.

For the Admin Center home-first path, use the focused Textual Pilot benchmark:

pytest -s -m slow tests/ace/tui/bench_admin_center_open.py

It reports # dispatch-to-paint p50/p95/max values for empty and populated project/config stores. Treat zero mounted concrete panes and comparable results across fixture sizes as the hard acceptance criteria; use the timing table for before/after comparison rather than adding a unit-test wall-clock threshold.

For the sase-zu Agents-tab load-tiering path, use the source/index oracle benchmark:

just bench-agent-load-tiering

It builds a temporary synthetic SASE_HOME with about 13,000 artifact directories by default, rebuilds the agent artifact index once, then compares authoritative source-scan rows with bounded and full-history index rows through the same Agents-tab query evaluator. The report includes p50/p95/max wall time, read/decode counters, decoded_records_per_returned_row on production paths, and speedups versus the source scan for source_scan, index_bounded, index_full_history, and the production loader's production_bounded and production_full_history paths, plus a periodic Tier 1 revalidate, a settled unchanged-query refresh session, and missing/extra row diffs. The synthetic fixture includes a production-shaped population of marker-only waiting/question records; production_bounded must stay at or below 3.0 decoded records per returned row while still filling the requested viewport. Pass --artifact-count, --runs, --warmup, --requested-limit, --session-refreshes, and repeated --query flags while iterating; --fixture-root reuses (or creates) the fixture under a directory, and --sase-home measures an existing archive instead (see Measured acceptance above).

Run via pytest:

pytest -s -m slow tests/perf/bench_tui_trace.py

Or as a script (writes a baseline numbers file the next phase can diff):

python -m tests.perf.bench_tui_trace --output ~/.sase/perf/tui_perf_baseline.json

The script also accepts explicit trace and key-to-paint output paths:

python -m tests.perf.bench_tui_trace \
  --output ~/.sase/perf/tui_perf_baseline.json \
  --trace-path ~/.sase/perf/tui_trace.jsonl \
  --perf-path ~/.sase/perf/tui_jk.jsonl

Fixture sizes:

Patches: 100,  500, 2000   (tests/perf/fixtures.py: PATCH_SIZES)
Agents:       50,  200, 1000   (tests/perf/fixtures.py: AGENT_SIZES)
Large reply:   1,    5,   20 MB (LARGE_REPLY_SIZES_MB)

Scenarios per fixture size:

  • cold start
  • query change
  • repeated query edits
  • 50-key j/k burst
  • auto-refresh with no changes
  • large-reply select

The per-scenario summary records wall-clock times, then aggregates p50 / p95 / max for every trace span and key-to-paint action observed during that scenario.

View-hints scenarios and committed baseline

The Agents-tab v keypath has its own scenario set, run separately because it needs disk-backed fixtures — the hint render reads raw_macros.md, *_prompt.md, and live_reply.md from a real artifacts dir:

large_reply_first_press           v on a plain agent with a 100 KB reply, cold header summary
large_reply_repeat_press          v again on the same row after the bar is torn down
session_container_press           v on a 5-member session container at the default metadata level;
                                  conversation content is full under the shared hint cap
session_container_unfolded_press  the same row at FoldLevel.FULLY_EXPANDED; foldable metadata grows,
                                  while conversation visibility and the shared hint cap stay unchanged
hint_mode_auto_refresh            an Agents-tab refresh tick while hint mode is active

Spans are sliced per step rather than pooled, so the plain-agent and session-container costs can be compared independently. Each step also carries a hint_counters block (annotated_chars, hints, commit_views, header_summary, agent_session_container) so a duration change can be attributed to the document actually getting smaller rather than to a quieter machine.

Run just these scenarios and print the table:

pytest -s -m slow tests/perf/bench_tui_trace.py::test_view_hints_scenario

Regenerate the committed baseline (5 runs, median per step and span, plus every raw run):

python -m tests.perf.bench_tui_trace --view-hints-baseline
# writes tests/perf/baselines/view_hints_baseline.json

Compare against tests/perf/baselines/view_hints_baseline.json rather than a transient capture. Two caveats when reading it:

  • wall_ms measures key dispatch through Pilot settle, so it carries unrelated repaint work and is much larger and much noisier than the spans. Compare the per-step spans table, not wall_ms.
  • An agents.view_hints_refresh span can appear inside a press step: a detail repaint often lands inside that step's settle window. That is real behavior, not bookkeeping noise.

Run the regression floor after changing this path:

just view-hints-perf-check

The floor compares traced spans against the committed baseline and ignores wall-clock Pilot settle time. It also checks that warm repeat presses and unchanged auto-refreshes do not rescan annotated text, and that session rows stay within the shared hint scan cap at both metadata levels. If long output is capped, sase's TUI shows a dim notice in the detail panel; hints are not generated past that notice. The committed baseline remains the synchronous pre-optimization reference and is not rewritten merely because conversation sections became fold-inert.

Prompt keys (<space> / <ctrl+n/p>, epic sase-1ex)

SASE_TUI_PERF=1 records key-to-paint samples for the prompt keys through the same JKPerfTimer (app._jk_perf) as the j/k harness, with one action per key and tab=current_tab:

  • prompt_space — begins at action_start_agent_from_patch entry; the model timestamp lands in activate() (reveal or fresh mount) when the bar focuses, and paint lands on the next refresh.
  • prompt_cycle_ctrl_p / prompt_cycle_ctrl_n — begin when the key reaches _handle_vcs_mru_cycle_key (after macro arg-name completion declines it); the model timestamp lands once the cycle edit applies, and paint lands on the next refresh.

Samples append to ~/.sase/perf/tui_jk.jsonl (SASE_TUI_PERF_PATH overrides); the disabled path is a true no-op and the JSONL path and env vars are unchanged from the j/k harness. Run the slow bench (prints one paint+handler table, asserts no budgets — shared-host timing is noisy):

pytest -s -m slow tests/ace/tui/bench_prompt_bar_keys.py

The bench seeds an isolated SASE home with about 30 MRU entries (five launchable projects in canonical plus display spellings, Patch refs, one stale entry per prune class) and covers first <space> after startup, steady <space>, single ctrl+p, a six-key ctrl+p burst at 50 ms cadence, and a first-visit ctrl+p; it also prints the stall-watchdog row count. The fast smoke test runs the smallest case end to end under just check:

pytest tests/ace/tui/test_prompt_key_perf_smoke.py

Later sase-1ex phases assert zeros through the I/O probe helper (tests/ace/tui/_prompt_key_io_probes.py), a context manager counting main-thread-only MRU reads/writes, list_project_records calls (facade plus by-name importers), subprocess.Popen constructions, ArtifactWatcher start/stop, and threading.Thread.join:

with prompt_key_io_probe() as probe:
    await pilot.press("ctrl+p")
    await pilot.pause()
probe.assert_quiet()

Final results (epic sase-1ex acceptance, bead sase-1ex.12)

Final rerun on 2026-10-03 with every phase landed, against the sase-1ex.1 baseline (shared host, so cross-run timing is noisy; paint_ms / handler_ms):

case baseline final verdict
first <space> (n=1) 321.02 / 280.75 101.93 / 74.99 improved, still above target
steady <space> p50 / p95 168.58 / 242.86 114.40 / 143.30 improved, p95 above target
single ctrl+p (n=1) 518.44 / 509.89 32.16 / 0.99 handler target met
burst ctrl+p p50 / p95 369.54 / 500.79 6.55 / 12.17 paint and handler targets met
first-visit ctrl+p (n=1) 451.68 / 442.31 16.24 / 0.61 spike gone, handler target met

Stall-watchdog rows during the final run: 1 (baseline: 0; host noise, not attributed).

Target verdicts: the synchronous handler is at most ~1 ms everywhere (target: at most 5 ms, met); burst key-to-paint p95 is 12.17 ms (target: at most 16 ms, met); the multi-hundred-millisecond first-press and first-visit spikes are gone. <space> p95 stays above the 60 ms target (143.30 ms steady, 101.93 ms first) even with the sase-1ex.11 hidden hot spare landed, so the remaining lever is the overlay-docked prompt bar experiment, recorded as a follow-up on the bead. A cold snapshot never blocks the first <space>: it opens a blank bar at once and late-prefills only an untouched session. No budget assertions live in the slow bench by design; the committed regression gates are the structural tests below, all cheap enough for just check:

  • warm ctrl+p / ctrl+n, first-visit getters, and warm <space> perform zero main-thread MRU reads/writes, list_project_records calls, Popens, watcher start/stop, and thread joins (test_launchable_mru.py, test_space_prefill.py, test_prompt_catalog.py, test_prompt_key_perf_smoke.py);
  • each prompt text-area mount/unmount/worker body runs once, base-first (test_prompt_mount_dedup.py);
  • a stale snapshot generation never overwrites text or reorders an in-flight cycle, and a cold <space> followed by typing is never clobbered by a late prefill (test_launchable_mru.py, test_space_prefill.py).

Pager bench

The pager benchmark (epic sase-1es) measures open, keystroke, search, memory, leak, and cold-start cost against a deterministic synthetic corpus over the 500 / 2k / 20k / 100k line ladder. Each case runs in a fresh subprocess with a per-case timeout (TIMEOUT is reported, never hung), plus a trivial-app floor so CPU numbers read both raw and above the floor:

just bench-pager
just bench-pager --cases code-sparse,log-dense --max-lines 2000 --no-cold
just bench-pager --cases markdown --ladder 500,2000 --output /tmp/pager-bench.json

The corpus generator is tests/perf/_pager_bench_corpus.py; the harness is tests/perf/bench_pager.py (slow). test_pager_bench_smoke.py is the fast non-slow smoke test that runs the smallest case end to end. Baseline numbers are not committed: shared-host timing is noisy, so record before/after tables in bead notes instead.

The committed tripwires are the regular tests, not the bench: dismissed-view release in tests/pager/test_view_leak.py, and the strip-cache bound plus the 20k-line tracemalloc ceiling in tests/pager/test_perf_gates.py.

Targets per phase gate

The targets below come from sdd/research/202604/sase_perf_research.md and are restated here so each phase agent has a single page to check against. A phase is green when the relevant targets are met without regressing any other span.

j/k highlight p95             < 16 ms
  (Stitches next/prev/up10: documented CommitsTimeline carve-out ≤ 30 ms;
  unmodified-master baseline was stitches.next 16.47 / stitches.up10 17.95,
  conform verification observed stitches.next 20.17 serial / 24.84 under xdist,
  agent-pane landing verification observed 26.59 serial / 27.22 under xdist)
key-to-paint p95              < 33 ms
debounced detail paint        < 150–250 ms
warm Patch reload, 1k    < 100 ms
no-change auto-refresh stall  ~0 ms (event-driven path; Phase 7)
large reply first paint       immediate plain render, syntax later/optional

Per-phase responsibilities:

  • Phase 2 (Patch j/k hot path): widget.patch_list.update_list call count drops to zero for j/k navigation; update_highlight p95 < 16 ms at 500 patches.
  • Phase 3 (data layer): warm Patch reload < 100 ms at 1k specs; patch.filter p95 should drop materially after the snapshot cache and query context land.
  • Phase 4 (agent panel + list): agents.refresh_panel_highlights and widget.agent_list.update_highlight p95 < 16 ms at 1k agents.
  • Phase 5 (incremental loader): agents.load_from_disk near zero on a no-change auto-refresh.
  • Phase 6 (artifact + render caching): widget.prompt_panel.update_display / widget.file_panel.update_display immediate first paint on the largest reply fixture.
  • Phase 7 (event-driven auto-refresh): no-change auto-refresh shows no agents/patch spans firing at all.

Adding a new span

from sase.ace.tui.util.trace import tui_trace

with tui_trace("module.name", count=len(items)):
    ...

Names use dotted lowercase. Counters should be ints / strs only — the emitter falls back to str(...) for unknown types via default=str, but keeping payloads JSON-friendly speeds downstream jq slicing.

When a span boundary forces a refactor (most existing hot paths split into foo() → _foo_impl() so the wrapping context manager doesn't fight indentation rules), keep both methods next to each other and let the public name stay the trace span name.

Reading the Admin Center Perf view

Perf is the eighth view in the Statistics tab of the Admin Center in sase's TUI. Open Admin Center with #, press 6 for Statistics, then press 0 followed by 8; [ / ] also cycle to it. The selected Statistics range applies, and g groups latency by subsystem, provider, or workflow. Perf is global rather than project-scoped: the project chip stays visible but is marked not applied because the underlying telemetry and TUI logs do not carry project attribution.

There is no CLI rendering of this dashboard. Use sase telemetry status for store and configuration state or sase telemetry health for the related traffic-light health assessment.

What the view shows

Five headline tiles summarize Startup, Stalls, Launch, Agent p95, and LLM p95. Startup, stalls, and launch timing come from bounded TUI diagnostic logs; agent and LLM latency come from the local telemetry store. Each tile shows one number, and they are deliberately not the same statistic:

Tile Value Detail line
Startup median visible-ready time sessions in range · slowest initial tab
Stalls stall count hitch count · worst freeze duration
Launch p95 total launch time launches · slow stages
Agent p95 p95 agent-run duration agent runs
LLM p95 p95 LLM invocation duration error rate · retry rate

The detailed body contains:

  1. Startup breakdown — p50, p95, and maximum durations for process→mount, mount→first paint, visible ready, and all surfaces ready, plus the slowest session in the range. The Startup tile itself reports the median visible-ready time and grades it OK below two seconds, warning from two to five seconds, and critical at five seconds or more.
  2. Stalls & hitches — per-event counts, worst and median duration, recency, suppressed counts, a Freeze records by context ranking, and recoveries. The two tiers are independent watchdogs with different thresholds (hitch at 1.5 seconds, stall at 5, both configurable — see Freeze and hitch capture), so one freeze long enough to trip the stall threshold has already tripped the lower hitch threshold and is recorded as both. The two counts therefore overlap and must not be added together. The tile reports stalls only and names hitches separately in its detail line. Any hitch makes the tile warn; any stall makes it critical.
  3. Latency & reliability — telemetry-backed p50, p95, maximum, a count, error rate, and retry rate. The count column is labeled by what it actually counts: LLM invocations under provider grouping, Agent runs under workflow grouping, and a plain Count under subsystem grouping. Provider grouping also shows input/output token and cache data, plus a note that provider rows without LLM invocation samples fall back to agent runs (which is why those rows read — for Err%/Retry%). Subsystem grouping omits the Share column entirely, because its rows measure different things and share no denominator; a subsystem with no counter renders — rather than 0. The configured telemetry.health_thresholds grade every row in this panel as well as the Agent p95 and LLM p95 tiles, using the same rules as sase telemetry health.
  4. Data & instrumentation — telemetry enablement, selected resolution, store size, raw/rollup counts, write freshness, and one coverage row per diagnostic log. Coverage reports file presence, records in the selected window, earliest retained record, truncation, and unreadable lines. A final line lists the optional probe flags; see Deep profiling and probe flags for how to read it.

The dashboard reads these files from ~/.sase/logs/ by default:

  • tui_startup.jsonl
  • tui_stalls.jsonl
  • tui_launch_timing.jsonl
  • tui_agent_loads.jsonl
  • tui_git_ops.jsonl
  • tui_external_tools.jsonl

Missing, disabled, truncated, or partially unreadable sources degrade independently, so the rest of the view remains usable.

The two data sources compute percentiles differently, and the in-app ? help names both methods:

  • Log-derived numbers (startup stages, launch timing, stall medians) use nearest-rank on the sorted sample, at index round(q * (n - 1)) clamped to [0, n - 1].
  • Telemetry-derived numbers (the Latency & reliability rows and the Agent p95 / LLM p95 tiles) are estimated by linear interpolation between cumulative histogram bucket bounds. They are bucket estimates, not exact sample percentiles.

Perf counts come from the telemetry store and the TUI logs, never from the agent-artifact index, so they are not comparable with the run counts on Overview, Projects, or Macros.

Every range except All time also loads the immediately preceding window of the same length; that second load is what the Startup and Agent p95 tiles compare against to show a delta. The other three tiles never show one.

Retention and rollups

The view's two data sources age out on entirely different rules:

  • Each TUI JSONL diagnostic log is byte-bounded, not time-retained: nothing expires on a clock, but the oldest records fall off once the file grows. The default limit is 2 MiB per current file, and rotation preserves exactly one .1 segment before discarding the previous one. Set SASE_TUI_TELEMETRY_MAX_BYTES to override the byte limit.
  • The telemetry store is time-retained and rolled up: raw samples for 48 hours, five-minute rollups for 30 days, and hourly rollups for 365 days by default. These durations are configurable under telemetry.retention.

So a long lookback reads telemetry at the coarser rollup resolution, and reads only the current plus rotated segment of each JSONL log. All time means everything still retained under those two rules, not an unbounded history.

Deep profiling and probe flags

Both probes are off by default and write to their own files under ~/.sase/perf/, which is a different directory from the ~/.sase/logs/ diagnostic logs the dashboard reads. The dashboard only names these flags; it never parses what they produce.

  • SASE_TUI_PERF=1 records per-keystroke j/k key-to-paint samples in ~/.sase/perf/tui_jk.jsonl by default. Override the path with SASE_TUI_PERF_PATH.
  • SASE_TUI_TRACE=1 records hot-path spans in ~/.sase/perf/tui_trace.jsonl by default. Override the path with SASE_TUI_TRACE_PATH; see Trace recorder.
  • SASE_TUI_HEAP=1 records periodic top heap allocation sites in ~/.sase/perf/tui_heap.jsonl by default. Override the path with SASE_TUI_HEAP_PATH; see Heap sampler.

Each probe records only when its variable is exactly 1. The Probes line in Data & instrumentation is looser: it prints on whenever the variable is set in the environment that started the TUI, so a deliberate SASE_TUI_PERF=0 still displays as on while nothing is being recorded. The same applies to SASE_TUI_TRACE and SASE_TUI_HEAP. Read that line as "set / unset", and check the value yourself when an expected probe file stays empty.

sase tui --tmux turns the perf and trace probes on unless the caller has already set the variable, so SASE_TUI_TRACE=0 sase tui --tmux … (or the SASE_TUI_PERF=0 equivalent) opts out. SASE_TUI_HEAP stays explicit because snapshots add overhead. Use just view-hints-perf-check for the automated hint-mode regression floor.

Bead history-independence gate (epic sase-1h8, phase sase-1h8.14)

The A1 criterion says hot-path bead latency must not move with closed history: on 1x to 8x scaled corpora, p95 of ready, default list, detail read, and a local mutation moves less than 10%, with warm point read <= 20 ms, active list <= 50 ms, and TUI board no-change refresh < 100 ms. The gate lives in tests/perf/bench_bead_scale.py --check-gate (unit-tested by tests/perf/test_bead_scale_gate.py). list is the paged active-status query that default sase bead list runs (list_issue_page), not the unbounded list_issues dump; the TUI op is the cached-snapshot (previous=) path.

# Strict local check: all criteria at <10%, 1x/2x/4x/8x.
just bead-perf-scale -- --check-gate
# CI gate: 1x+4x, runner-noise tolerance, known misses reported not skipped.
just bead-perf-scale-gate

CI (just bead-perf-scale-gate, tolerance 0.5) blocks only on the ready p95 ratio. The four ratio criteria that still miss, plus all three absolute ceilings, run as recorded known-misses (--gate-allow): shared-runner wall clocks are contention-sensitive, and the misses below are real scaling gaps with follow-ups, not noise. Drop ids off the --gate-allow list as follow-ups land; the strict local run keeps enforcing everything.

Before/after (medians, ms)

Before is the bench-phase baseline (bead sase-1h8.1 notes); after is the sdd/plans/202610/perf_artifacts/bead_perf_gate.json sweep (runs=5, core pin 5c4033f6, host load ~27, so small-op medians are inflated — clean-window probes in parentheses).

op 1x before 1x after 8x before 8x after
detail read 495 37.8 (6.9) ~4,800 60.3
ready 408 34.9 (4.1) 4,100 38.5
default list (paged) 775† 65.8 (25.8) 8,200† 318.1
note append 511 132.7 (19.4) — 257.0
update 549 134.6 (19.1) — 258.7
TUI no-change refresh 33 15.6 (9.8) — 133.2
remote-backed CLI note 4,446 ~1,500 (1,075) — —
audited sase bead read (live store) 5,000–8,000 3,045 — —

† Before measured the unbounded list_issues; after measures the paged default-list query. Unbounded list still costs ~456 ms at 1x (Python hydration of ~6,900 rows).

Verdict and known misses

ratio:ready passes (1.03 at 1x to 8x): ready no longer replays closed history. The rest miss and stay open with measured breakdowns (each miss is owned by a task bead):

  • ratio:list (5.1x) and abs:active-list (318 ms at 8x): the paged query touches only active rows, but hydration cost is per row and the active set itself grows 8x with the corpus. Needs bounded serving for unbounded active lists. Owned by sase-1iw.
  • ratio:detail (marginal formally, 3.7x in clean probes) and abs:point-read (25 ms at 4x vs 20 ms): the bead_show_issue_detail binding itself scales (~6 ms at 1x to ~23.5 ms at 4x) while plain bead_show is flat (0.6 to 0.8 ms), so the relations expansion scans corpus-sized state. Facade hydration adds ~nothing. Owned by sase-1iv. The cause is neighborhood_in in sase-core read_model/queries.rs, which loads every issue id and scans every link_provenance row.
  • ratio:note (1.8x) and ratio:update (2.1x; binding-level 2.8x/2.9x in the epic notes): per-mutation cost still grows with stream count. An strace count during one facade append (1,383 stat calls total including interpreter startup, under 2,000 streams) shows no full per-mutation stat sweep on that path, so the sweep theory needs revisiting — the follow-up owns the breakdown into admission, publication, and binding overhead. Owned by sase-1iu.
  • abs:tui-nochange (133 ms at 8x, doubling per scale doubling): the no-change path cost is linear in active rows. Needs row virtualization. Owned by sase-1ix.

Contention caveat: this shared host regularly sits at load 20+. One formal sweep was discarded outright (1x ready p95 88 ms vs 4 ms minutes later on the same store). Ratios are measured within a single run so both scales share the window, every report records /proc/loadavg, and borderline verdicts should be re-run before acting on them.