sase tui Performance Runbook¶
This runbook explains how to capture and compare performance data for sase's TUI, the
sase tui terminal user interface. It started as the Phase 1 deliverable for the TUI
performance overhaul (bead sase-w.1, sdd/epics/202604/tui_perf_overhaul_1.md), and
later performance phases still rely on the tracing and benchmark harness described here.
Athena host-relief baseline¶
Bead sase-zn.1 captured this host-pressure baseline on athena on 2026-09-11 while
investigating prompt-input lag in a long-lived sase tui session. Use it as the
reference point for the later sase-zn phases that measure remaining CPU, heap,
refresh, and scratch pressure.
| signal | before | after host relief |
|---|---|---|
/tmp tmpfs |
20G used, 13G available, 62% full | 5.6G used, 26G available, 18% full |
/ filesystem |
837G used, 29G available, 97% full | 815G used, 51G available, 95% full |
| swap | 30.4G used of 64G | 13.1G used of 64G |
| sase's TUI process | PID 2019865, 10036968 kB RSS, 3332948 kB swap | PID 2351038, 766372 kB RSS, 0 kB swap |
| SASE scratch matches | 24 matched /tmp/*cargo-target*, /tmp/*core-target*, /tmp/sase-*-recovery*, and synced managed-root build targets; 37G total |
0 remaining matches |
SASE_TMPDIR |
/home/bryan/tmp/sase, via /home/bryan/tmp -> Sync/home/tmp |
/home/bryan/.cache/sase/tmp in chezmoi-managed .profile and the live sase's TUI process |
The cleanup removed only SASE-named cargo/core target directories, SASE recovery
bundles, and the stale ~/.sase/perf/tui_trace.jsonl file. The durable environment
change is in the chezmoi source home/dot_profile, applied to ~/.profile, then sase's
TUI was restarted in tmux pane sase:5.1 after sourcing the updated profile.
Py-spy profiles from the same run are stored outside the repo at:
~/.sase/perf/athena-host-relief-before-20260911T1542.svg~/.sase/perf/athena-host-relief-after-20260911T1546.svg
Both profile runs reported delayed sampling under remaining host load, so treat the flamegraphs as attribution hints rather than precise wall-clock percentages. The resource deltas above are the reliable phase-1 baseline.
Athena phase-8 verification¶
Bead sase-zn.8 remeasured the live athena session on 2026-09-12 at 13:15-13:20 EDT,
while the host was under real agent load: load average 38.93 / 37.40 / 37.07, many
run_agent_runner.py processes were active, and pytest workers from several SASE
workspaces were consuming CPU.
| signal | 2026-09-12 verification |
|---|---|
/tmp tmpfs |
22M used of 32G; no /tmp/*cargo-target*, /tmp/*core-target*, or /tmp/sase-*-recovery* matches remained |
/ filesystem |
847G used, 19G available, 98% full |
| swap | 8.3G used of 64G |
| sase's TUI process | PID 54331, up 5h09m; VmRSS: 1199148 kB, VmSwap: 275464 kB, Anonymous: 1163148 kB |
| sase's TUI CPU | main thread sampled at 569 ticks over 60s, about 0.095 cores, or roughly 2.3 CPU-hours/day; whole process was about 0.30 cores under load |
| SASE scratch | live SASE_TMPDIR=/home/bryan/.cache/sase/tmp; managed cargo-targets/ was present and bounded at 68K |
| old hot frames | 25s py-spy raw capture had no reconcile_agent_artifact_index_dismissed_family_members samples and only 4/2399 read_current_notifications_snapshot samples |
| watchdog | not green: the last 30 minutes contained 17 tui_hitch and 5 tui_pump_hitch starts, with fresh Agents-tab hitches during the run |
| key-to-paint | not re-confirmed green: the live sase's TUI process was not started with SASE_TUI_PERF=1; the existing tui_jk.jsonl was stale from 2026-09-10 and showed Agents p95 174.06 ms |
The profile is stored outside the repo at
~/.sase/perf/athena-verify-sase-zn8-20260912T131736.raw. py-spy reported sampling
delays and 550 read errors under host load, so use the capture for top-frame attribution
only. The phase-8 result is therefore mixed: memory pressure, scratch placement,
main-thread CPU, and the original reconcile/notification hot frames look bounded, but
the watchdog and key-to-paint evidence do not prove the responsiveness target is green.
The remaining hitches point at Agents-tab refresh/render work rather than the original
sase-zn hot sites.
Agents query load tiering¶
Bead sase-zu.7 removed the sase-zu load-tiering beta flags on 2026-09-13 and
remeasured the not machine:apollo startup cliff. The key behavior is now
unconditional: safe indexed query pushdown stays on Tier 1, unsupported queries defer
full-history reconciliation instead of blocking first paint, and full-history artifact
index loads use the index before falling back to a source scan.
Reproduce the synthetic archive benchmark with the workspace virtualenv:
.venv/bin/python tests/perf/bench_agent_load_tiering.py \
--output ~/.sase/perf/agent_load_tiering_sase-zu.7_20260913.json
The 2026-09-13 run used sase-core-rs 0.34.24 and a 13,000-artifact fixture with
not machine:apollo, 5 measured runs, 1 warmup run, and requested_limit=100.
| path / signal | before evidence | after evidence |
|---|---|---|
| startup first paint | 2026-09-12 04:57 EDT Tier 1 indexed load: 688 index rows, agents_ready_seconds=8.036 |
Real archive direct loader from this checkout: Tier 1 artifact-index load, 820 records, 100 returned rows, no error |
| source-scan cliff | 2026-09-12 05:42, 05:56, 06:20 EDT Tier 2 source scans: agents_ready_seconds=23.336, 24.395, 42.663 |
Synthetic source scan remains the expensive baseline: p50 8621.35 ms, p95 8891.94 ms |
| bounded indexed first window | Regressed path escalated to full source scan for not machine:apollo |
Synthetic Tier 1 indexed bounded path: p50 107.17 ms, p95 113.48 ms, 113 visible rows |
| full-history indexed parity | Source scan was the only complete-history path for this query | Synthetic full-history index path: p50 7672.78 ms, p95 7986.75 ms, missing_count=0, visible_extra_count=0 |
| live log tail | 2026-09-13 tui_agent_loads.jsonl still contains old-process full reloads before this checkout is deployed |
2026-09-13 tui_startup.jsonl tail shows Tier 1 artifact-index startups, latest agents_ready_seconds=8.013 |
Interpretation: the measured first-window load is again in line with the indexed 04:57
baseline rather than the 06:20 source-scan cliff. The bounded window is intentionally
incomplete (has_more=true) while the complete indexed path proves parity with the
source scan. If a future query cannot be narrowed by the artifact index, the loader
should report the visible result as incomplete and defer history; it should not increase
the startup read to the full archive.
Measured acceptance (sase-zu.8.5)¶
Bead sase-zu.8.5 remeasured the repaired tree on 2026-09-13 15:00-15:30 EDT: main tree
d698f92e05, sase-core-revision.txt 23f19f0 built by just install-venv from the
linked checkout (sase-core-rs 0.34.24 dev build, index schema 30, scan wire 9). The
sase-zu.7 full-history rows above measured a cached index read, not production's
revalidating path, and are superseded by these numbers. The harness now routes every
path, the source scan reference included, through the tab's
compute_apply_loaded_agents dismissal step, reports read/repair/decode counters, a
periodic Tier 1 revalidate, and a settled unchanged-query session:
.venv/bin/python tests/perf/bench_agent_load_tiering.py \
--fixture-root ~/.cache/sase/tmp/agent-load-tiering-13k --session-refreshes 10 \
--output ~/.sase/perf/agent_load_tiering_sase-zu.8.5_synthetic13k_20260913.json
.venv/bin/python tests/perf/bench_agent_load_tiering.py --sase-home ~/.sase \
--output ~/.sase/perf/agent_load_tiering_sase-zu.8.5_athena_real_20260913.json
Both runs used not machine:apollo, 5 measured runs, 1 warmup and requested_limit=100
on a host running other agents. Speedups divide the source-scan p50 by the path's p50;
the session baseline models the pre-epic behavior of one source scan per broad refresh.
13,000-artifact synthetic fixture (warm page cache, 270 hidden rows):
| path | p50 / p95 / max ms | read / repair / decode work | speedup | missing / extra |
|---|---|---|---|---|
| source scan | 8933.0 / 9711.8 / 9711.8 | 13,000 dirs, 26,003 marker files parsed | 1.00x | reference |
| bounded first paint | 101.1 / 110.8 / 110.8 | 0 marker checks, 104 records decoded | 88.32x | bounded prefix |
| actual full history | 9280.9 / 9928.0 / 9928.0 | 12,938 marker checks, 0 repaired, 12,803 decoded | 0.96x | 0 / 0 |
| periodic Tier 1 revalidate | 136.0 / 174.9 / 174.9 | 339 marker checks, 0 repaired, 204 decoded | - | - |
| 10-refresh settled session | 10130.5 total | 1 full-history read, 12,938 checks, 13,947 decoded | 10.6x | - |
Athena real archive (11,183 source artifacts):
| path | p50 / p95 / max ms | read / repair / decode work | speedup | missing / extra |
|---|---|---|---|---|
| source scan | 10866.4 / 11253.6 / 11253.6 | 11,183 dirs, 41,272 marker files parsed | 1.00x | reference |
| bounded first paint | 1002.4 / 1202.8 / 1202.8 | 0 marker checks, 1,062 records decoded | 10.84x | bounded prefix |
| actual full history | 2988.8 / 4381.5 / 4381.5 | 7,201 marker checks, 0 repaired, 1,389 decoded | 3.64x | 13 / 1 |
| periodic Tier 1 revalidate | 1087.6 / 1300.6 / 1300.6 | 6,813 marker checks, 0 repaired, 1,273 decoded | - | - |
| 10-refresh settled session | 10315.0 total | 1 full-history read, 7,201 checks, 13,071 decoded | 12.6x | - |
Stage attribution. On the warm synthetic fixture the raw Rust calls are close: source
scan p50 1916 ms, cached full-history index read 1624 ms, revalidating full-history read
2283 ms. Revalidation's source-directory reconcile plus 12,938 marker signatures is the
epic-added cost (about 660 ms). The remaining time is per-row Python work that both
paths share: dict-to-wire conversion (about 1.3 s), Agent decode (about 2.0 s) and the
agents-live filter building one field entry per row (about 2.9 s). A single
non-selective full-history load is therefore not faster than the source scan on a warm
fixture, and this benchmark does not claim it is. On the real archive the raw Rust
source scan is 5924 ms against a 1029 ms revalidating index read, because the index
drops hidden, dismissed and non-matching rows before decode (956 of 11,183 records).
Most of the Tier 1 revalidate cost there is hidden-row repair: about 6,800 marker
checks.
Real-archive parity. All 13 missing rows are --code and --mon members of dismissed
sessions. The Rust index hides descendants of a dismissed session root (sase-core
34b3229, before this epic), but the tab's Python dismissal step, which the source scan
reference uses, does not. The one extra row was an agent created during the run. Before
the dismissal step was added, 107 rows differed; every one was a dismissed identity or
one of those session members.
Live session: sase's TUI from this tree (sase tui -x --tab agents, PID 2078791, logs
in ~/.sase/perf/sase-zu.8.5-tui_{startup,agent_loads,trace}.jsonl) against the real
archive, with the not machine:apollo query committed interactively:
| signal | observation |
|---|---|
| startup | Tier 1 artifact-index startup, on_mount_to_first_paint_seconds=0.414, agents_ready_seconds=6.457 |
| query commit | bounded filter load, 1506 ms, 171 rows, has_more=true |
| automatic history upgrade | exactly one input_quiet_tier2_reconcile, 7154 ms, complete_history=true |
| ordinary refreshes, 14 min | 13 cached index auto_refresh loads (p50 1421 ms, max 3181 ms), no repeated full-history load |
| periodic revalidate | 3 tier1_index_revalidate loads, p50 1570 ms |
| exact deltas | 20 watcher, starting-poll and notification deltas, p50 about 150 ms |
| explicit Full history | ,y, then f in the Refresh panel: one manual_full_history load, 4066 ms, complete_history=true |
| remaining expensive stage | 3 bounded loads before the upgrade fell back to a 6.7-7.4 s source scan: artifact index operation lock busy |
First paint stays in the indexed baseline range (8.036 s on 2026-09-12; 5.5-6.8 s in the
live sase's TUI today). Unchanged covered history is loaded once per committed query.
The lock-busy fallbacks come from the process-local index lock's 50 ms read timeout,
which 62ad9b657c added. Old artifact directories are outside the bounded startup
inotify watch budget. Their marker changes therefore reach the tab through revalidation,
not watcher deltas: a touched July done.json showed up in the index as a new
done_sig without a rebuild.
Idle-host CPU diet¶
An idle sase host (TUI open, routines running, no agent work) used to burn roughly four
cores: job subprocesses at ~109 spawns/min, each paying ~0.6s of import tax, plus the
TUI's unconditional full reconcile every refresh_interval. The idle-CPU-diet epic
(sase-wn) turns ticks cheap instead of rare. Use this recipe when a host looks busy at
rest again — it should take about five minutes.
Success criteria (idle, no agent work):
| signal | budget |
|---|---|
| sase sustained CPU | < 0.4 cores |
| job subprocess spawns | < 25/min across every routine |
| representative job import | < 0.2s / < 400 modules |
| idle TUI | < 10% of a core, ~0 stall-watchdog records / 30 min |
| idle axe collector file opens | near zero when nothing changed |
The spawn budget sits on a floor no fs trigger can remove: stale_running_cleanup's
real input is process liveness (a dead PID touches no file), so it stays on the always
trigger in both the hooks (5s) and checks (300s) lanes — ~12/min at full idle on its
own. The rest of the shipped-config floor is max_quiet: 120s re-fires from the nine
fs-guarded jobs (~4.5/min), the run_every: 30s waits-lane pair (~4/min), and the ≥60s
always-trigger lanes (~2/min): ≈ 22/min total. At the post-diet import cost (< 0.2s per
boot) that floor is well under a tenth of a core, which is why the budget is a low spawn
rate, not zero; a periodic proc-liveness cache could shave the stale_running_cleanup
term but has not been worth new machinery.
Deterministic regression floors live in tests, not wall-clock CI:
- job-SDK import closure:
tests/test_chop_import_budget.py(andtests/test_idle_cpu_diet_guardrails.py) - zero-spawn idle tick for fs-guarded hooks-lane jobs:
tests/test_axe_default_chop_triggers.py/tests/test_idle_cpu_diet_guardrails.py
Five-minute diagnosis¶
# 1. Fleet spawn rate and no-op ratio (reads routine metrics.json)
sase axe routine status
# 2. Confirm the human overlay matches on-disk counters
jq '{chops_spawned, chops_no_op, chops_skipped, last_tick_spawns, last_tick_skipped,
spawn_rate_per_minute, no_op_ratio}' \
~/.sase/axe/lumberjacks/*/metrics.json
# 3. Process CPU over a quiet minute (/proc deltas)
pidof -x sase; ps -o pid,pcpu,pmem,comm -p $(pgrep -d, -f 'sase (tui|axe)')
# sample twice, 60s apart:
awk '{print $1,$14,$15}' /proc/$(pgrep -n -f 'sase tui')/stat; sleep 60; \
awk '{print $1,$14,$15}' /proc/$(pgrep -n -f 'sase tui')/stat
# utime+stime ticks / (HZ * elapsed) is cores used. On Linux HZ is usually 100.
# 4. Stall-watchdog count over a 30-minute idle TUI session
jq -s 'map(select(.event=="tui_stall" or .event=="tui_hitch")) | length' \
~/.sase/logs/tui_stalls.jsonl
# 5. Per-tick TUI counters (surfaces reloaded, axe collector file opens)
SASE_TUI_TRACE=1 sase tui
# quit after a few idle refresh intervals, then:
jq -c 'select(.span=="refresh.auto_tick")' ~/.sase/perf/tui_trace.jsonl | tail
jq -c 'select(.event=="axe.collect")' ~/.sase/perf/tui_trace.jsonl | tail
# 6. Job import cost (wall-clock; the test suite locks the module set, not time)
python -X importtime -c "import sase.jobs.sdk" 2>&1 | tail -1
python -c "import sase.jobs.sdk, sys; print(len(sys.modules))"
# 7. Routine run-history gap analysis (idle ticks should be skipped, not no_op)
# History lives under the legacy lumberjacks/<routine>/chops/<job>/ paths.
python - <<'PY'
from pathlib import Path
import json
root = Path.home() / ".sase" / "axe" / "lumberjacks"
for routine in sorted(root.iterdir()):
jobs_dir = routine / "chops"
for job in sorted(jobs_dir.glob("*")) if jobs_dir.exists() else []:
index = job / "index.json"
if not index.exists():
continue
ids = json.loads(index.read_text())[:8]
print(f"== {routine.name}/{job.name} ==")
for run_id in ids:
meta = json.loads((job / "runs" / f"{run_id}.json").read_text())
print(f" {meta.get('started_at')} {meta.get('status')} {meta.get('reason')}")
PY
A quiet host after the diet should show sase axe routine status Job load at or
below the ~22/min shipped-config floor — fs-guarded jobs contribute 0.0 between
max_quiet re-fires and their lanes report last tick 0 spawned, while ~12/min of the
floor is stale_running_cleanup and is expected — and refresh.auto_tick records with
surfaces_reloaded=0 (or only axe/notifications when those tokens actually moved)
and axe_file_opens near zero. If spawn rate is well above the floor, check whether
shipped jobs lost their fs trigger in src/sase/default_config.yml, or whether
ace_refresh_tokens is off.
Suite test-cost gate¶
The repository-wide pytest cost harness is separate from sase's TUI trace spans. Use it when a change may affect the whole test suite's cost model rather than one TUI interaction:
just test-cost
The lane runs the default fast suite, records per-file wall/CPU time, collection time,
the worker RSS curve, and attributed hot causes (now with a per-cause CPU-seconds
breakdown alongside wall seconds and count), then prints tools/test_cost_report.
Recordings live under ${SASE_HOME:-~/.sase}/test-selection/<project>/timings/cost/;
set SASE_TEST_COST_DIR to redirect them. tools/check_test_cost_budgets compares the
newest recording with tests/perf/baselines/test_cost_budgets.json, and
just check-full plus CI Telemetry's test-cost job enforce that comparison.
Severity model: hard vs. advisory¶
A single host regularly runs several agents' test suites at once, and Python's
time.perf_counter() wall-clock durations stretch under CPU contention that has nothing
to do with the suite's own cost -- a causes.ace_page_enter wall time can rise 17%
between a quiet host and a busy one for the exact same test run (same invocation count).
CPU seconds, invocation counts, and RSS do not move with contention, so the gate is
split by metric family instead of failing everything uniformly:
| family | metrics | default | tolerance |
|---|---|---|---|
| wall | causes.<name> (limit), total_file_wall_seconds, idle_seconds, collection_seconds |
advisory | ci/local |
| cpu | causes.<name>.cpu (cpu_limit), total_file_cpu_seconds, collection_cpu_seconds |
hard | cpu |
| count | causes.<name>.count (count_limit) |
hard | none |
| memory | peak_worker_rss_kib, median_worker_rss_kib, post_collection_worker_rss_kib |
hard | ci/local |
Advisory failures are printed with full detail but never fail the gate and are
excluded from the exit code -- they are the "the host was probably busy" bucket.
Hard failures fail the gate; they are the contention-stable metrics, so an overage
there is a real signal. Any budget entry can override its dimension's family default
with an explicit enforce (for the wall limit), cpu_enforce (for cpu_limit), or
count_enforce (for count_limit) key set to "hard" or "advisory". --suggest
always writes these explicitly so the committed file never silently depends on the
implicit default.
Count budgets are compared without the cpu/ci/local tolerance: a count_limit
already carries ~25% headroom over the worst observed invocation count
(round_up_nice(max_observed_count * 1.25)), because counts are near-deterministic --
the headroom belongs in the limit itself, not in a runtime tolerance. That policy trips
on a refactor that doubles a call site but tolerates a handful of newly added tests.
CPU attribution note: CostRecorder.measure() times time.process_time(), which is
process-wide, so CPU burned by other coroutines on the same event loop during an awaited
span is attributed to that span. This matches how wall time already behaves, and within
one xdist worker tests run sequentially, so the attribution stays meaningful.
Sleep-dominated causes such as pilot_pause_delay report near-zero CPU seconds -- that
is correct, and is exactly why the count dimension matters for them.
Run the gate with --strict to make advisories fatal too, for deliberate investigation
on a quiet host:
tools/check_test_cost_budgets --strict
just test-cost runs the fast suite and then tools/check_test_cost_budgets
(non-strict, so a wall-only advisory no longer breaks the lane). Inside
just check-full, tools/run_silent discards a wrapped command's captured output on
success, which would otherwise hide a printed advisory; check-full follows the wrapped
just test-cost line with an unwrapped
tools/check_test_cost_budgets --report-advisories, which re-reads the recording
test-cost just wrote and always exits 0, so advisories still reach the operator.
Budget entries are either per-worker or suite-wide totals, and mixing the two up silently defeats the gate:
collection_secondsandcollection_cpu_secondsare per-worker limits: every xdist worker collects the whole suite, sobuild_cost_record()sums the metric across workers before a budget entry marked"per_worker": truedivides it back down by the record's worker count (falling back to the number of worker payloads that reported collection time, then to 1) before comparing against the limit.- Every other summary entry (
total_file_wall_seconds,total_file_cpu_seconds,idle_seconds,peak_worker_rss_kib,median_worker_rss_kib, andpost_collection_worker_rss_kib) and everycauses.*entry is a suite-wide metric, not normalized by worker count. - Each worker records
start,post_collection,median, andpeakRSS summaries. The suite-levelworker_rss_curve_kibtakes the maximum worker value forstart,post_collection, andpeak; itsmedianis the median of every positive worker summary value across those four fields, and itssample_countis the sum of worker sample counts. It is an aggregate summary, not one process's time series. peak_worker_rss_kib,median_worker_rss_kib, andpost_collection_worker_rss_kibare flat aliases for the corresponding suite-level curve fields. The report and budget suggestion tools do not divide these RSS values by worker count.
Read failures by bucket:
total_file_wall_secondsandidle_secondspoint to broad suite cost or waiting, but are advisory-only -- pair them withtotal_file_cpu_seconds(hard) to tell contention from a real regression.collection_seconds,collection_cpu_seconds, and post-collection RSS point to import-time state. Median and peak RSS summarize later retained or run-time growth, but the suite median is over worker summary fields rather than raw time-series samples.- Cause entries such as
textual_app_run_test_enter,ace_page_enter,parser_create,yaml_load, andsubprocess_runpoint to the hot pattern to audit; each reports up to three failures (causes.<name>wall/advisory,causes.<name>.cpuhard,causes.<name>.counthard).
For focused diagnosis, rerun the cost lane on a path or node and print more rows:
just test-cost -- tests/ace/tui/widgets/test_vim_normal_key_containment.py
tools/test_cost_report --top 20
The committed budgets intentionally have tolerance for host noise. Raise a limit with this workflow, not a hand-picked number, and do not raise one to hide a one-off regression:
just test-cost
tools/check_test_cost_budgets --suggest
--suggest reads the newest retained recordings (--history N to limit the sample),
derives each wall/cpu limit as ceil(worst recorded value / (1 + tolerance)) (using the
local/ci tolerance for wall metrics and the cpu tolerance for CPU metrics) and
each count_limit as ceil(max observed count * 1.25) with no tolerance, rounds each
up to a round number, and prints a budget JSON -- with
enforce/cpu_enforce/count_enforce written explicitly on every entry -- along with
the notes provenance line (sample size, host, UTC date range, worker-count range,
node-count range, and per-metric min/median/max) to paste alongside the new limits.
tools/check_test_cost_budgets --ci defaults to
os.environ.get("GITHUB_ACTIONS") == "true", not a bare CI variable: an agent runtime
that exports CI=true locally must not silently switch a local run onto the wider CI
tolerance. GitHub Actions always sets GITHUB_ACTIONS; nothing else does. Pass --ci
explicitly to force the CI tolerance on a non-GitHub-Actions host.
Trace recorder¶
SASE_TUI_TRACE=1 enables tui_trace(...) context managers spread across the Patch,
agents, and AXE hot paths. Each entered span emits one JSONL line to:
~/.sase/perf/tui_trace.jsonl
Override the destination with SASE_TUI_TRACE_PATH=/tmp/foo.jsonl. When the env flag is
unset the context managers are near-zero-cost no-ops.
Each record contains at least:
ts unix epoch seconds
span dotted span name (e.g. "agents.refresh_panel_widgets")
duration_ms wall time inside the span
current_tab "artifacts" | "agents" | "services" | null
…plus any per-call counters (count, agents, panels, output_bytes, …) and any
global context fields seeded via sase.ace.tui.util.trace.set_trace_context(...) (the
app pushes current_tab and current_idx automatically).
Point-in-time records emitted by trace_event(...) contain event instead of
span/duration_ms. They are used for selection and highlight watcher transitions
where there is no timed block to measure, and for the prompt-submit pipeline:
launch.accepted fires in the handler that unmounts the prompt bar, and
launch.submitted fires when the durable sase run proc takes over, carrying
accept_to_submit_ms and the pending-launch stages it waited in. Pair them by
launch_id
(jq -c 'select(.event|startswith("launch."))' ~/.sase/perf/tui_trace.jsonl) to measure
how long a submit stayed pending after the bar disappeared.
Timed spans for the main Patch, agents, and AXE hot paths (by file, relative to
src/sase/ace/tui/; run rg 'tui_trace\(' src/sase for the complete set, which also
covers loader stages and pager opens):
actions/patch/_display.py—patch.refresh_display,patch.refresh_debounced,patch.refresh_detail_onlyactions/patch/_loading.py—patch.filteractions/agents/_display.py—agents.refresh_display,agents.refresh_display_incremental,agents.refresh_debouncedactions/agents/_display_panel_widgets.py—agents.refresh_panel_widgetsactions/agents/_display_panel_layout.py—agents.refresh_panel_highlights,agents.refresh_focused_panelactions/agents/_loading_helpers.py—agents.load_from_diskactions/agents/_loading_live_hints.py—agents.live_hint_refreshactions/agents/_display_detail_render.py—agents.view_hints_refreshactions/hints/_files.py—agents.view_files,agents.view_agent_files,agents.view_hint_bar_mountwidgets/prompt_panel/_agent_display_hints.py—widget.prompt_panel.update_display_with_hintswidgets/patch_list.py—widget.patch_list.update_list,widget.patch_list.update_highlight,widget.patch_list.patch_patch_rowwidgets/patch_detail.py—widget.patch_detail.update_displaywidgets/agent_list.py—widget.agent_list.update_list,widget.agent_list.update_highlight,widget.agent_list.patch_agent_row,widget.agent_list.try_remove_rowswidgets/agent_detail.py—widget.agent_detail.update_display,widget.agent_detail.update_display_immediatewidgets/artifacts/relation_panel.py—widget.relation_panel.update_relationswidgets/prompt_panel/_agent_display.py—widget.prompt_panel.update_display,widget.prompt_panel.update_header_only,widget.prompt_panel.update_tribe_display,widget.prompt_panel.refresh_slow_tool_metadata_from_cachewidgets/prompt_panel/_agent_display_header_summary.py—widget.prompt_panel.build_detail_header_summaryand one child span per resolver (see "SASE CONTEXT enrichment" below)widgets/file_panel/_panel.py—widget.file_panel.update_displaywidgets/llm_calls_panel.py—widget.llm_calls_panel.update_displaywidgets/axe_dashboard.py—widget.axe_dashboard.update_display,widget.axe_dashboard.update_lumberjack_overview(routine overview),widget.axe_dashboard.update_chop_run_display(job detail)
Spans nest cleanly: a single keypress that fires agents.refresh_debounced will record
one outer span plus inner widget.agent_list.update_highlight and
agents.refresh_panel_highlights spans.
Three agents-tab spans carry counters meant for soak assertions on tribe-panel stability:
agents.finalize_query_filterreportsagents_in(the filter's input roster) andagents_out(its output). The old singleagentscounter was the input, which is easy to misread as the published roster size.agents.apply_loaded_agents_preparedreportsfinalize_planasapplied,discarded, orabsent. A discard also carriesfinalize_plan_discard_reason(stale_tokenorroster_fingerprint). A soak that never seesappliedproves nothing about the off-thread finalize path: every apply fell back to the inline one.agents.refresh_panel_widgetsreportspanel_widget_ids, the tribe-stable widget ids in mount order, so a tribe panel's continuous presence is assertable from the trace.
Agents-tab paint frames (agents.paint_frame)¶
Counters such as the ones above pass on an idle tab, so they cannot show what a panel
looks like while a node joins it. With SASE_TUI_TRACE=1 every completed Agents
display refresh (and every out-of-band change of the agent-list column width) also emits
one agents.paint_frame event: kind (full_rebuild, incremental, highlight,
container_width), the refresh's source / display_cost / fallback_reason, the
app's app_grouping_mode, the selected_identity, #agent-list-container's negotiated
container_width (its styles.width) beside the laid-out container_layout_width, and
a panels list. Each panel entry carries widget_id, object_id (the widget's Python
id(), so a remount is visible), option_count, collapsed, collapse_intent,
grouping_mode, requested_width, height, highlighted, highlighted_identity,
scroll_y and viewport_height. seq orders the frames.
# One row per frame: what changed on screen between consecutive refreshes.
jq -c 'select(.event == "agents.paint_frame")
| {seq, kind, source, display_cost, fallback_reason, container_width,
panels: [.panels[] | [.widget_id, .option_count, .requested_width, .height]]}' \
~/.sase/perf/tui_trace.jsonl
A node joining @epic should produce exactly one frame whose panels or
container_width differ from the previous one. A full_rebuild frame carrying
fallback_reason: stale_grouping_mode, or a container_width frame right after a
refresh frame that already widened a panel's requested_width, is the flicker itself.
tests/ace/tui/test_epic_panel_arrival_frames.py asserts these invariants
deterministically; the paint log lives in actions/agents/_paint_log.py.
The arrival's display_cost says which path painted it. An ordinary node (no clan,
session or workflow relation, no wider than the panel's existing rows, no new banner)
records display_row_insert: the row is inserted in place and no update_list runs. A
node that cannot be inserted still rebuilds only its own panel, recording
display_panel_rebuild after a display_row_insert record whose fallback_reason
names the gate that declined it (width_growth, status_membership_change,
workflow_tree_change, panel_membership_change). A STARTING agent is not rendered, so
the apply that first shows it is the arrival, not a status-bucket move. Tally the paths
over a soak:
jq -r 'select(.event == "agents.paint_frame" and .kind != "settled")
| [.display_cost, .fallback_reason] | @tsv' ~/.sase/perf/tui_trace.jsonl \
| sort | uniq -c | sort -rn
A soak that never creates a node shows no display_row_insert at all; that proves
nothing about arrivals.
Rebuild scope is decided per panel. When a roster change concerns only some panels (a
duplicated identity, a BY_STATUS bucket change, or a change to a panel's workflow
tree), the apply rebuilds those panels and patches the rest: display_panel_rebuild on
an incremental frame, never display_full_rebuild. Each panel it names is one
agents.refresh_work event with stage: display_fallback, the fallback_reason
(panel_membership_change, status_membership_change, workflow_tree_change or
group_fold_change), and panel, that panel's widget id. A group_fold_change rebuild
means only a grouping-banner fold changed: the roster is identical, but the panel's rows
were built under different fold state. Tally which panels a soak rebuilt, and why:
jq -r 'select(.event == "agents.refresh_work" and .stage == "display_fallback"
and .panel != null)
| [.panel, .fallback_reason] | @tsv' ~/.sase/perf/tui_trace.jsonl \
| sort | uniq -c | sort -rn
A display_full_rebuild therefore means the whole tab was rebuilt: a changed search
query, an unsupported or stale grouping mode, or a fast path that refused the apply.
To look at the same arrival by eye, capture the @epic panel before and after a node
joins it with a live sase screenshot. This is a human check and stays out of the
golden lane (live captures carry real timestamps and host state):
# Before: capture a fresh TUI on the Agents tab and keep its tmux window.
sase screenshot --keep -o /tmp/epic_before.png -w "@epic" -- -t agents
# Copy the printed sase_tmux_target=..., launch a cheap agent that lands in the epic
# tribe (a prompt tagged `#tribe:epic`), wait for its row to appear, then:
sase screenshot --window <sase_tmux_target> -o /tmp/epic_after.png
The tmux-launched TUI that sase screenshot starts gets SASE_TUI_TRACE=1 injected
automatically, so its frame events land in ~/.sase/perf/tui_trace.jsonl to compare.
Compare the two PNGs for the panel border, scroll position, highlighted row and column
width, and confirm the same arrival in the agents.paint_frame events above. Do not
copy either PNG under tests/ace/tui/visual/snapshots/png/; goldens are maintained only
by just fix-tui-screenshots.
sase's TUI deliberately keeps live-workspace pencil hints off the startup-critical
agents loader. The first load classifies only cheap persisted diff_path badges. After
that agents list has applied, agents.live_hint_refresh runs VCS probes for active,
non-terminal rows that do not yet have a persisted diff and patches changed rows in
place. During startup investigations, treat agents.load_from_disk and
agents.live_hint_refresh as separate costs: the former controls time to first
interactive Agents-tab paint, while the latter explains deferred pencil-badge updates.
Reading a view-hints capture¶
Pressing v on the Agents tab nests four spans, outermost first:
agents.view_files whole v keypath (both tabs)
└─ agents.view_agent_files Agents-tab branch only
├─ widget.prompt_panel.update_display_with_hints the annotated render
└─ agents.view_hint_bar_mount mounting the HintInputBar
agents.view_files is the keypress → hint-bar-mounted interval: today the bar is
mounted last, so this span is effectively the render cost plus the mount cost.
Subtracting agents.view_hint_bar_mount from it gives the part of the wait the user
pays for work that is not the bar appearing.
agents.view_hints_refresh is the same annotated render fired again from an Agents-tab
detail repaint or from the detail-header enrichment message, rather than from a
keypress. Seeing it repeatedly with hint mode active is the signal that the document is
being rebuilt on refresh.
Useful counters:
agents.view_files tab
agents.view_agent_files agent_session_container, hints, commit_views,
header_enrichment_pending, outcome
(mounted | refocused | empty | no_agent |
detached_container)
agents.view_hints_refresh agent_session_container, hints, commit_views
update_display_with_hints agent_session_container, hints, commit_views,
tool_call_reports, annotated_chars,
header_summary (warm | cold)
annotated_chars counts every character handed to the hint scanner, summed across
fragments and session members, so it is the size term to divide a duration by.
header_summary says whether the render had a warm detail-header summary: a cold
render omits the SASE CONTEXT hints entirely and will be rebuilt when the enrichment
worker lands.
Slice one press out of a capture with:
jq -c 'select(.span | startswith("agents.view_") or . == "widget.prompt_panel.update_display_with_hints")
| {span, duration_ms, hints, annotated_chars, header_summary, agent_session_container, outcome}' \
~/.sase/perf/tui_trace.jsonl | tail -20
Trace events currently wired include:
selection.current_idx.setwidget.patch_list.watch_highlightedand.suppressedwidget.agent_list.watch_highlightedand.suppressedwidget.bgcmd_list.watch_highlightedand.suppressed
SASE CONTEXT enrichment (detail-header summary)¶
build_detail_header_summary (widgets/prompt_panel/_agent_display_header_summary.py)
is the worker that resolves the PLAN, BEAD, ARTIFACTS, MEMORY, SKILLS, and
WORKSPACES lanes of the SASE CONTEXT section (bead sase-l6.1,
plans/202608/sase_context_incremental.md). It emits one parent span,
widget.prompt_panel.build_detail_header_summary, plus one child span per resolver, all
sharing that dotted prefix:
widget.prompt_panel.build_detail_header_summary (parent)
.macros_used
.bead_display
.plan_enrichment
.slow_tool_sources
.agent_page_url
.linked_delta_groups
.artifact_file_paths the one resolver with no cache — usually the
most expensive lane before phase `stores` lands
.artifact_reads
.memory_reads
.skill_uses
.opened_workspaces
.delta_entries
.wait_bead_statuses
The parent span carries agent (the agent's cl_name) and cache_state ("cold" on
the first time this process has resolved that agent identity, "warm" after). This
marker is process-local and best-effort — it tracks whether this worker has touched
the identity before, independent of and coarser than each resolver's own on-disk cache —
so it exists to make a raw capture readable without cross-referencing every resolver's
cache state.
Since phase stream (bead sase-l6.4), one selection's enrichment worker resolves its
requested lanes cheapest-first across LANE_RESOLUTION_BATCHES
(_agent_display_header_summary.py) and merges/publishes each batch as it lands, so a
single selection now emits up to three
widget.prompt_panel.build_detail_header_summary parent spans back to back — one per
non-empty batch — instead of one. All three share the same agent; only the first
carries cache_state: "cold" for a never-before-seen identity; the batches after it
reuse the same process-local seen-set and read "warm". Group by agent and read the
spans in capture order to see which lane group landed first (typically the
free/cached-lookup batch) versus last (typically the store-backed
ARTIFACTS/MEMORY/SKILLS batch, before phase stores's caches are warm).
Slice one selection's enrichment out of a capture with:
jq -c 'select(.span | startswith("widget.prompt_panel.build_detail_header_summary"))
| {span, duration_ms, agent, cache_state}' \
~/.sase/perf/tui_trace.jsonl | tail -20
Reproduce the baseline table (per-resolver cold/warm cost, plus where
artifact_file_paths spends its time inside list_artifact_files) with the committed
benchmark, which drives the exact same spans read above rather than re-timing the
resolvers by hand:
pytest -s -m slow tests/perf/bench_detail_header_summary.py
python -m tests.perf.bench_detail_header_summary --include-home \
--count 20 --output ~/.sase/perf/detail_header_summary_baseline.json
--include-home is required to measure real ~/.sase agents; without it the script
only exercises a tiny hermetic in-memory fixture, so it stays safe to run in CI.
--count controls how many non-clan agents from
load_tiered_agents(full_history=False) are sampled.
Heap sampler¶
SASE_TUI_HEAP=1 enables an opt-in tracemalloc sampler for long-lived sase's TUI
sessions. Tracing starts when AceApp initializes, while each snapshot is scheduled
from the TUI timer into a pump-free task and written from a worker thread. Samples
append one compact JSONL record to:
~/.sase/perf/tui_heap.jsonl
Override the destination with SASE_TUI_HEAP_PATH=/tmp/tui_heap.jsonl. The default
interval is 300 seconds; override it with SASE_TUI_HEAP_INTERVAL_SECONDS. One sample
is also written at startup. Each record includes current_bytes, peak_bytes, and the
top allocation sites grouped by source line, each with its captured traceback. The
sampler records 25 sites with 25 traceback frames by default; tune them with
SASE_TUI_HEAP_TOP_N and SASE_TUI_HEAP_NFRAME. It keeps no previous snapshots in
memory, so use adjacent JSONL records to compare growth over time.
Quick capture¶
SASE_TUI_TRACE=1 sase tui
# … exercise the path you care about (cold start, query change, j/k burst,
# auto-refresh, large reply select) …
# Quit with q.
# Inspect:
jq -c 'select(.span | startswith("widget.agent_list."))' \
~/.sase/perf/tui_trace.jsonl | head -20
To inspect point events instead of timed spans:
jq -c 'select(.event)' ~/.sase/perf/tui_trace.jsonl | head -20
For key-to-paint timing during j/k navigation, also enable the separate perf recorder:
SASE_TUI_TRACE=1 SASE_TUI_PERF=1 sase tui
jq -c . ~/.sase/perf/tui_jk.jsonl | head -20
Override the key-to-paint path with SASE_TUI_PERF_PATH=/tmp/tui_jk.jsonl.
Agents that launch the TUI via sase tui --tmux get SASE_TUI_TRACE=1 and
SASE_TUI_PERF=1 injected automatically; export the variable to 0 before invoking to
opt out. SASE_TUI_HEAP is never auto-enabled; set it explicitly for heap attribution.
Offline Fleet fault j/k benchmark¶
Use the committed offline Fleet fault benchmark when changing unified remote projection, federation refresh, remote lifecycle row metadata, or Agents navigation around remote rows:
just test-slow tests/ace/tui/bench_tui_jk.py -k fleet_jk_fault_scenarios
The benchmark uses in-memory federation fixtures only: no enrolled machines, network
access, or real federation worker are involved. It warms the rendered remote rows, then
records Agents-tab j/k key-to-paint samples while each fault script is active:
hung_hostkeeps a healthy host's rows while a second host contributes a deadline diagnostic.reconnect_churnkeeps rows visible while one host reports aging/reconnecting health.event_burstapplies a larger multi-host catalog burst.
Read the printed p50/p95/max tables per scenario. Every next and prev p95 must be
< 16 ms, and the run must leave no tui_stalls.jsonl rows. A p95 failure points at
the Agents remote-row key-to-paint path; any stall row means the event loop or Textual
pump was blocked and must be investigated before landing.
Unread and idle baseline (epic sase-1d7)¶
Use the committed unread benches when changing unread acknowledgment, unread navigation, the 1 Hz runtime tick, or fleet reprojection. They record the baseline the later epic phases compare against:
just test-slow tests/ace/tui/bench_tui_jk.py -k "unread_bulk_ack or unread_jump"
just test-slow tests/perf/bench_tui_trace.py -k idle_tick_and_fleet_refresh
Both benches drive a screenshot-shaped roster (about 200 agents, clan containers, three
tribes) with in-memory fixtures only: the notification dismiss write is stubbed, so no
real notifications are touched. The unread bench covers ,u and ,j with the target
visible, inside a collapsed panel, inside a collapsed clan, and on another Agents tab;
the idle bench times one countdown tick (_patch_agent_runtime_rows) and one
fleet_refresh reprojection (forced and unchanged variants) by direct call, because the
tick and fleet apply are skipped while navigating and the j/k benches cannot see them.
The committed ceilings are deliberately generous (order-of-magnitude regression only);
compare the printed p50/p95/max tables against the bead-note baseline instead.
The same spans can be captured live from a real session:
SASE_TUI_TRACE=1 SASE_TUI_PERF=1 sase tui
# ... press ,u / ,j /,J on the Agents tab, then quit with q ...
jq -c 'select(.span == "leader.unread_bulk_ack" or .span == "leader.unread_jump"
or .span == "unread.ack_complete" or .span == "unread.reconcile"
or .span == "unread.chrome_apply"
or .span == "agent_nodes.projection_index")' ~/.sase/perf/tui_trace.jsonl
jq -c 'select(.action == ",u" or .action == ",j" or .action == ",J")' \
~/.sase/perf/tui_jk.jsonl
Note: an off-tab ,j that switches Agents tabs records its paint under the
agents_tab_switch action, not ,j — the switch starts its own sample. The
store_bytes counter on unread.ack_complete is statted on the ack worker thread; the
UI thread never stats the store itself.
Freeze and hitch capture¶
sase's TUI always-on watchdog writes event-loop and Textual message-pump diagnostics to
~/.sase/logs/tui_stalls.jsonl. There are two independent severity tiers:
tui_hitch/tui_pump_hitchfire after 1.5 seconds by default. These are compact records containing the current tab and selection, the last action, keypress age, and the main-thread stack. Recovery rows use the corresponding*_recoveredevent and include the episode duration. Hitch records deliberately omit full asyncio-task and worker thread dumps, and are rate-limited;suppressed_countreports episodes omitted since the previous admitted record.tui_stall/tui_pump_stallretain the existing 5-second threshold and richer task/thread diagnostics. A long freeze can produce both a hitch and a stall because the two state machines are independent.
Lower both tiers for a short verification soak with:
SASE_TUI_HITCH_THRESHOLD_SECONDS=0.25 \
SASE_TUI_PUMP_HITCH_THRESHOLD_SECONDS=0.25 \
SASE_TUI_STALL_THRESHOLD_SECONDS=0.75 \
SASE_TUI_PUMP_STALL_THRESHOLD_SECONDS=0.75 \
SASE_TUI_STALL_POLL_INTERVAL=0.02 \
SASE_TUI_PUMP_STALL_POLL_INTERVAL=0.02 \
SASE_TUI_STALL_PATH=/tmp/sase-tui-soak.jsonl \
sase tui
Exercise startup typing, launch and cleanup bursts, prompt-history and revive-agent modal opens, tab switching during agent churn, and refreshes while the artifact or dismissed-bundle indexes are contended. A fixed path should stay interactive and produce no hitch/stall rows. Inspect any records in chronological order:
jq -c '{event, stall_seconds, duration_seconds, current_tab, current_idx,
last_action, last_keypress_age_s, suppressed_count, main_thread_stack}' \
/tmp/sase-tui-soak.jsonl
Read main_thread_stack from its final frame upward to find the blocking call. For a
pump-only event, the loop may still be running while one Textual message pump is
awaiting slow work; a loop event means the event loop itself stopped servicing the
watchdog beacon. Disable the compact tiers independently with SASE_TUI_HITCH_DISABLE=1
and SASE_TUI_PUMP_HITCH_DISABLE=1; the existing stall-tier disable flags remain
SASE_TUI_STALL_DISABLE=1 and SASE_TUI_PUMP_STALL_DISABLE=1.
Hitch and stall rows (and their recoveries) carry additive whole-process fields:
lateanddetected_by—true/"watchdog_lateness"when the watchdog's own poll was late even though the loop beacon had already run (the GIL-race blind spot: a stop-the-world pause on another thread freezes every thread, and the beacon recovers first). On-time rows reportfalse/"loop_gap"(or"pump_gap").poll_lag_s— the poll lateness minus one poll interval (0.0on on-time rows).net_stall_seconds— the duration minus one poll interval, floored at the tier threshold.stall_seconds/duration_secondsare unchanged.app_instance_id— the per-instance ID minted at app construction, so a busy hour spanning restarts splits cleanly per instance.- Recovery rows add
gc_overlap_s(GC pause seconds overlapping the episode),gc_generations, andgc_triggers, attributed from the GC telemetry ring. A hitch caused by an intentional idle collection is then distinguishable from an interactive freeze.
Filter one instance's late freezes with:
jq -c 'select(.event == "tui_hitch" and .late == true) | {ts, stall_seconds,
poll_lag_s, net_stall_seconds, gc_overlap_s, app_instance_id}' \
~/.sase/logs/tui_stalls.jsonl
The persistent TUI diagnostic JSONL files under ~/.sase/logs/—stall, git-operation,
launch-timing, external-tool, agent-load, and startup records—rotate independently
before appending a record would make a non-empty file exceed 2 MiB. Each keeps one .1
generation. Set SASE_TUI_TELEMETRY_MAX_BYTES to another per-file byte limit, or 0
for no size rotation. This bound is separate from the opt-in trace files under
~/.sase/perf/.
GC pause and memory heartbeat rows¶
The GC telemetry service (src/sase/ace/tui/util/gc_telemetry.py, installed from
_start_post_first_paint_services) adds two more row kinds to tui_stalls.jsonl:
tui_gc_pause— one row per gen-2 collection and per collection ≥ 50 ms, withgeneration,duration_s,thread,trigger("automatic"unless the collection ran undergc_trigger(...), e.g. an intentional idle collection),collected/uncollectable, andapp_instance_id. Rows are rate-capped at 60/min;suppressed_countreports pauses folded into a row since the previous admitted one. Heartbeat totals stay exact either way.tui_memory_heartbeat— every 5 minutes (plus once ~15 s after startup), withrss_bytes/vmswap_bytes(from/proc/self/status,nulloff Linux),major_faultsandmajor_faults_delta(from/proc/self/stat),gc_count/gc_threshold/gc_freeze_count, exact per-generationcount/total_s/max_sfor the window they cover (gc_generations), instanceuptime_s(since install), andwindow_s(wall-clock span since the previous heartbeat, so GC share is exact astotal_s / window_s).
Every row carries app_instance_id (also on tui_startup and tui_agent_load rows),
so a busy hour spanning restarts splits cleanly per instance. Filter one instance's GC
share with:
jq -c 'select(.event == "tui_gc_pause" and .app_instance_id == "<id>")
| {generation, duration_s, trigger, thread}' ~/.sase/logs/tui_stalls.jsonl
Disable with SASE_TUI_GC_TELEMETRY_DISABLE=1. The service never auto-installs under
the sase.ace.testing harness.
The watchdog contributes exact hitch totals to each heartbeat through a stall_watchdog
heartbeat provider: loop_hitch_episodes / loop_hitch_seconds and
pump_hitch_episodes / pump_hitch_seconds count every recovered episode in the
window, while loop_suppressed_episodes / loop_suppressed_seconds (and the pump
equivalents) count the rate-limited ones. Row counts stay bounded; the totals stay
exact.
GC idle-collection policy¶
The idle-GC policy (src/sase/ace/tui/util/gc_policy.py, installed from
_start_post_first_paint_services right after telemetry) keeps automatic gen-2
collections off the interactive path:
- Once startup loads settle (
_mount_state_loads_done) and input first goes idle, onegc.collect()followed bygc.freeze()runs under thestartup_freezetrigger, exactly once per app instance. It is never re-frozen: cyclic garbage among frozen objects is never reclaimed. - On the classic three-generation collector the policy preserves the current
t0/t1and raisesthreshold2to 10_000, so automatic full collections become a rare backstop (about once an hour at 2.9 gen-1/s). Any other collector shape leaves thresholds alone and recordsthreshold_policy: "unsupported"(instead of"raised") on thetui_memory_heartbeatrow; the idle collector still runs. - A 1 s tick runs
gc.collect()on the UI thread only when input has been quiet for 3 s, no prompt is active, the navigation gate is idle, no modal screen is accepting text input, at least 60 s have passed since the last full collection of any trigger, and a collection is due (100 gen-1 collections since the last full one). Mouse clicks and scrolls count as input for the quiet window without consuming the events. - After 5 minutes without a full collection, or when RSS grew more than 500 MB since the
last full one (from the heartbeat's cached sample, never a UI-thread
/procread), the quiet window relaxes to 1 s. These run taggedbackstop/rss_backstop; the raisedthreshold2remains the hard backstop. - After an idle or backstop collection, freed arenas are released with
malloc_trimon a daemon worker thread (sase-tui-gc-trim, via thetrim_allocatorhalf ofsase.axe.runner_idle_memory), at most every 10 minutes.
Confirm it from tui_gc_pause rows: intentional collections carry trigger values
startup_freeze / idle / backstop / rss_backstop, while anything still tagged
"automatic" started on its own. An idle collection that overlaps a hitch is
attributable through the same trigger. Disable freeze, threshold, and the idle collector
with SASE_TUI_GC_POLICY_DISABLE=1 (telemetry stays on). The policy never auto-installs
under the sase.ace.testing harness, and teardown restores the original thresholds,
calls gc.unfreeze(), and stops the tick timer.
Frozen-share report¶
tools/tui_freeze_report is the read-only measuring tool for the freeze epic. It
reports per app_instance_id (rows that predate instance IDs are bucketed into
tui_startup windows instead of being dropped) over a --since / --until window:
tools/tui_freeze_report --since 2026-10-01T12:00 --until 2026-10-01T13:00
tools/tui_freeze_report --instance abc123 --path /tmp/sase-tui-soak.jsonl
Each instance reports the loop-only union frozen share (pump and loop tiers are never
double-counted), hitch duration median / p90 / max, the late versus on-time split, GC
share by generation and trigger, idle-collection pause p50 / p95 / max, automatic gen-2
collections within 2 s of input (a documented heuristic: it needs a
last_keypress_age_s context row within 2 s of the pause), and the RSS / swap
trajectory from heartbeats.
Startup telemetry capture¶
~/.sase/logs/tui_startup.jsonl (sase/logs/tui_telemetry.py:log_tui_startup) gets one
durable record per sase's TUI session, written after both the Agents and Services
surfaces finish their first load — the same point the visible startup-stopwatch badge
stops. It exists so every "startup dropped from X to Y" claim in this repo's plans and
epics is checkable against a real terminal run instead of a modelled component sum
(plans/202608/ace_startup_critical_path.md).
Each record carries two headline metrics, both measured from the App's on_mount — the
same anchor the visible stopwatch badge uses, so the two numbers should visually track
what you saw on screen:
all_surfaces_ready_seconds— elapsed time until both the Agents and Services tabs' first load finished, regardless of which tab was visible. This is today's stopwatch semantics.visible_ready_seconds— elapsed time until the initially visible tab's own surface was interactive. Recorded from day one even though nothing currently drives the stopwatch off of it, so a later change to end the stopwatch on the visible surface instead of both surfaces has an honest before/after rather than a redefinition that silently invalidates every prior capture.
The record also carries process_start_to_on_mount_seconds and
on_mount_to_first_paint_seconds (the two stages before either surface starts loading),
agents_ready_seconds / axe_ready_seconds (each surface's own elapsed time from
process start), initial_tab, agent_row_count, index_row_count (the Tier-1 index
query's row count when the load went through the persistent index, null otherwise),
and source / tier / artifact_source from the load's AgentLoadState.
process_start_to_on_mount_seconds keeps its historical meaning (the AceApp init stamp
to on_mount). Four additive fields split the wider OS-process-start → compose path
without changing that field:
interpreter_cli_import_seconds— OS process start to just beforeAceAppimportapp_module_import_seconds—from sase.ace.tui import AceAppapp_construct_seconds—AceApp()constructioncompose_seconds— Textual consumption ofcompose()
Missing stamps record as null (unit tests and non-CLI constructions).
Attributed startup capture (sase-132.1)¶
One command sequence reproduces a fully attributed capture at a recorded SHA. Store
outputs under ~/.sase/perf/ with the epic prefix and label the host as quiet or
busy from load average plus run_agent_runner.py count (see the script's
host_state).
# 1. Standalone loader bench against the real archive (label host state)
.venv/bin/python tests/perf/capture_tui_startup.py bench \
--sase-home ~/.sase \
--output-dir ~/.sase/perf
# 2. Three live traced startups + one pyinstrument profile
.venv/bin/python tests/perf/capture_tui_startup.py live \
--runs 3 --profile \
--output-dir ~/.sase/perf
# 3. Import-time of the TUI app module in this environment
.venv/bin/python tests/perf/capture_tui_startup.py importtime \
--output-dir ~/.sase/perf
Or all three:
.venv/bin/python tests/perf/capture_tui_startup.py all --output-dir ~/.sase/perf
Query startup-window contention without reconstructing the window from timestamps:
jq -c 'select(.startup_window==true)' ~/.sase/perf/sase-132.1-*/tui_trace.jsonl
Loader wait-vs-work inside the startup agents.load_from_disk span:
jq -c 'select(.span=="agents.load_from_disk"
or .span=="agents.load_from_disk.dismissed_snapshot"
or .span=="agents.load_from_disk.provider"
or .span=="agents.load_from_disk.index"
or .span=="agents.load_from_disk.decode"
or .span=="agents.load_from_disk.projections")' \
~/.sase/perf/sase-132.1-*/tui_trace.jsonl
Axe first load (duration span plus the legacy point event, both with file_opens):
jq -c 'select(.span=="axe.startup" or .span=="axe.load_status"
or .span=="axe.collect" or .event=="axe.collect")' \
~/.sase/perf/sase-132.1-*/tui_trace.jsonl
To capture a before/after pair:
# baseline, before your change
git stash # or check out the prior commit in a second workspace
sase tui # quit once the tabs finish loading; repeat 3x for a stable read
git stash pop
# after your change
sase tui # quit once the tabs finish loading; repeat 3x
Then compare the two sets of records:
jq -c '{timestamp, all_surfaces_ready_seconds, visible_ready_seconds,
agents_ready_seconds, axe_ready_seconds, agent_row_count,
index_row_count, source, tier, artifact_source,
process_start_to_on_mount_seconds,
interpreter_cli_import_seconds, app_module_import_seconds,
app_construct_seconds, compose_seconds}' \
~/.sase/logs/tui_startup.jsonl
tui_agent_loads.jsonl's slow-stage records are censored below
_DEFAULT_SLOW_LOADER_STAGE_THRESHOLD_SECONDS (2.0 s in
src/sase/ace/tui/actions/agents/_loading_disk_support.py); a capture run against a
tree that has already dropped under 2 s needs the sub-threshold stages too. Override the
threshold for the run instead of editing the constant (non-numeric or non-positive
values fall back to the default):
SASE_TUI_LOADER_LOG_THRESHOLD_SECONDS=0.05 sase tui
TUI import-budget ratchet (sase-13p)¶
tests/ace/tui/test_app_import_budget.py guards the fresh-interpreter module count for
import sase.ace.tui.app with a strict < against _MAX_MODULE_COUNT (measured plus
20). The cap only moves down: raising it requires a named module that is genuinely
needed on the first-paint path with deferral impossible, recorded in the test comment
with its commit, and the raise covers only the attributed amount. Never redefine success
as <= and never raise the cap to absorb unattributed drift. New TUI code must not add
eager imports to the startup closure; use use-site or TYPE_CHECKING imports and extend
the test's deferred_modules probe with the deferred heavy modules.
Measure and attribute with the closure tool (run it with the project venv):
.venv/bin/python tools/tui_import_closure
.venv/bin/python tools/tui_import_closure --diff $(git merge-base HEAD origin/master)
The bare form prints the current closure count. --diff REV measures the same import at
REV from a temporary detached worktree (removed afterwards) and lists the sase.*
modules added and removed.
Synthetic-data benchmark harness¶
The harness lives at tests/perf/bench_tui_trace.py. It generates in-memory Patch and
agent fixtures, then drives the TUI through Textual Pilot without touching real
~/.sase data. It is marked pytest.mark.slow, so it does not run as part of
just test.
For the Updates > Plugins catalog-scale path, use the focused benches and the committed measuring-stick baseline:
just bench-plugin-catalog-scale
That records p50/p95/max at 10 / 250 / 1000 / 2000 entries for pane open, one filter
keystroke, 20 j presses (through the queued OptionHighlighted handler), ' jump-hint
allocation, and one I mark toggle, plus non-TUI enrich/fetch cost curves. The
committed numbers live in tests/perf/baselines/plugin_catalog_scale_baseline.json.
Filter-keystroke and j-press p95 must stay under 16 ms at n=2000; eager enrich fetches
must stay O(installed) and enrich's scan_work (catalog rows walked, counted through
CountingEntries) must stay linear in n. Rewrite the recorded rows after a capture
run with python -m tests.perf.bench_plugin_catalog_scale --write-baseline and
SASE_PLUGIN_CATALOG_SCALE_WRITE_BASELINE=1 pytest -s -m slow tests/ace/tui/bench_plugins_catalog_scale.py.
just plugin-catalog-scale-check is the CI regression floor.
For the Admin Center home-first path, use the focused Textual Pilot benchmark:
pytest -s -m slow tests/ace/tui/bench_admin_center_open.py
It reports # dispatch-to-paint p50/p95/max values for empty and populated
project/config stores. Treat zero mounted concrete panes and comparable results across
fixture sizes as the hard acceptance criteria; use the timing table for before/after
comparison rather than adding a unit-test wall-clock threshold.
For the sase-zu Agents-tab load-tiering path, use the source/index oracle benchmark:
just bench-agent-load-tiering
It builds a temporary synthetic SASE_HOME with about 13,000 artifact directories by
default, rebuilds the agent artifact index once, then compares authoritative source-scan
rows with bounded and full-history index rows through the same Agents-tab query
evaluator. The report includes p50/p95/max wall time, read/decode counters,
decoded_records_per_returned_row on production paths, and speedups versus the source
scan for source_scan, index_bounded, index_full_history, and the production
loader's production_bounded and production_full_history paths, plus a periodic Tier
1 revalidate, a settled unchanged-query refresh session, and missing/extra row diffs.
The synthetic fixture includes a production-shaped population of marker-only
waiting/question records; production_bounded must stay at or below 3.0 decoded records
per returned row while still filling the requested viewport. Pass --artifact-count,
--runs, --warmup, --requested-limit, --session-refreshes, and repeated --query
flags while iterating; --fixture-root reuses (or creates) the fixture under a
directory, and --sase-home measures an existing archive instead (see
Measured acceptance above).
Run via pytest:
pytest -s -m slow tests/perf/bench_tui_trace.py
Or as a script (writes a baseline numbers file the next phase can diff):
python -m tests.perf.bench_tui_trace --output ~/.sase/perf/tui_perf_baseline.json
The script also accepts explicit trace and key-to-paint output paths:
python -m tests.perf.bench_tui_trace \
--output ~/.sase/perf/tui_perf_baseline.json \
--trace-path ~/.sase/perf/tui_trace.jsonl \
--perf-path ~/.sase/perf/tui_jk.jsonl
Fixture sizes:
Patches: 100, 500, 2000 (tests/perf/fixtures.py: PATCH_SIZES)
Agents: 50, 200, 1000 (tests/perf/fixtures.py: AGENT_SIZES)
Large reply: 1, 5, 20 MB (LARGE_REPLY_SIZES_MB)
Scenarios per fixture size:
- cold start
- query change
- repeated query edits
- 50-key j/k burst
- auto-refresh with no changes
- large-reply select
The per-scenario summary records wall-clock times, then aggregates p50 / p95 / max for every trace span and key-to-paint action observed during that scenario.
View-hints scenarios and committed baseline¶
The Agents-tab v keypath has its own scenario set, run separately because it needs
disk-backed fixtures — the hint render reads raw_macros.md, *_prompt.md, and
live_reply.md from a real artifacts dir:
large_reply_first_press v on a plain agent with a 100 KB reply, cold header summary
large_reply_repeat_press v again on the same row after the bar is torn down
session_container_press v on a 5-member session container at the default metadata level;
conversation content is full under the shared hint cap
session_container_unfolded_press the same row at FoldLevel.FULLY_EXPANDED; foldable metadata grows,
while conversation visibility and the shared hint cap stay unchanged
hint_mode_auto_refresh an Agents-tab refresh tick while hint mode is active
Spans are sliced per step rather than pooled, so the plain-agent and session-container
costs can be compared independently. Each step also carries a hint_counters block
(annotated_chars, hints, commit_views, header_summary,
agent_session_container) so a duration change can be attributed to the document
actually getting smaller rather than to a quieter machine.
Run just these scenarios and print the table:
pytest -s -m slow tests/perf/bench_tui_trace.py::test_view_hints_scenario
Regenerate the committed baseline (5 runs, median per step and span, plus every raw run):
python -m tests.perf.bench_tui_trace --view-hints-baseline
# writes tests/perf/baselines/view_hints_baseline.json
Compare against tests/perf/baselines/view_hints_baseline.json rather than a transient
capture. Two caveats when reading it:
wall_msmeasures key dispatch through Pilot settle, so it carries unrelated repaint work and is much larger and much noisier than the spans. Compare the per-stepspanstable, notwall_ms.- An
agents.view_hints_refreshspan can appear inside a press step: a detail repaint often lands inside that step's settle window. That is real behavior, not bookkeeping noise.
Run the regression floor after changing this path:
just view-hints-perf-check
The floor compares traced spans against the committed baseline and ignores wall-clock Pilot settle time. It also checks that warm repeat presses and unchanged auto-refreshes do not rescan annotated text, and that session rows stay within the shared hint scan cap at both metadata levels. If long output is capped, sase's TUI shows a dim notice in the detail panel; hints are not generated past that notice. The committed baseline remains the synchronous pre-optimization reference and is not rewritten merely because conversation sections became fold-inert.
Prompt keys (<space> / <ctrl+n/p>, epic sase-1ex)¶
SASE_TUI_PERF=1 records key-to-paint samples for the prompt keys through the same
JKPerfTimer (app._jk_perf) as the j/k harness, with one action per key and
tab=current_tab:
prompt_space— begins ataction_start_agent_from_patchentry; the model timestamp lands inactivate()(reveal or fresh mount) when the bar focuses, and paint lands on the next refresh.prompt_cycle_ctrl_p/prompt_cycle_ctrl_n— begin when the key reaches_handle_vcs_mru_cycle_key(after macro arg-name completion declines it); the model timestamp lands once the cycle edit applies, and paint lands on the next refresh.
Samples append to ~/.sase/perf/tui_jk.jsonl (SASE_TUI_PERF_PATH overrides); the
disabled path is a true no-op and the JSONL path and env vars are unchanged from the j/k
harness. Run the slow bench (prints one paint+handler table, asserts no budgets —
shared-host timing is noisy):
pytest -s -m slow tests/ace/tui/bench_prompt_bar_keys.py
The bench seeds an isolated SASE home with about 30 MRU entries (five launchable
projects in canonical plus display spellings, Patch refs, one stale entry per prune
class) and covers first <space> after startup, steady <space>, single ctrl+p, a
six-key ctrl+p burst at 50 ms cadence, and a first-visit ctrl+p; it also prints the
stall-watchdog row count. The fast smoke test runs the smallest case end to end under
just check:
pytest tests/ace/tui/test_prompt_key_perf_smoke.py
Later sase-1ex phases assert zeros through the I/O probe helper
(tests/ace/tui/_prompt_key_io_probes.py), a context manager counting main-thread-only
MRU reads/writes, list_project_records calls (facade plus by-name importers),
subprocess.Popen constructions, ArtifactWatcher start/stop, and
threading.Thread.join:
with prompt_key_io_probe() as probe:
await pilot.press("ctrl+p")
await pilot.pause()
probe.assert_quiet()
Final results (epic sase-1ex acceptance, bead sase-1ex.12)¶
Final rerun on 2026-10-03 with every phase landed, against the sase-1ex.1 baseline
(shared host, so cross-run timing is noisy; paint_ms / handler_ms):
| case | baseline | final | verdict |
|---|---|---|---|
first <space> (n=1) |
321.02 / 280.75 | 101.93 / 74.99 | improved, still above target |
steady <space> p50 / p95 |
168.58 / 242.86 | 114.40 / 143.30 | improved, p95 above target |
single ctrl+p (n=1) |
518.44 / 509.89 | 32.16 / 0.99 | handler target met |
burst ctrl+p p50 / p95 |
369.54 / 500.79 | 6.55 / 12.17 | paint and handler targets met |
first-visit ctrl+p (n=1) |
451.68 / 442.31 | 16.24 / 0.61 | spike gone, handler target met |
Stall-watchdog rows during the final run: 1 (baseline: 0; host noise, not attributed).
Target verdicts: the synchronous handler is at most ~1 ms everywhere (target: at most 5
ms, met); burst key-to-paint p95 is 12.17 ms (target: at most 16 ms, met); the
multi-hundred-millisecond first-press and first-visit spikes are gone. <space> p95
stays above the 60 ms target (143.30 ms steady, 101.93 ms first) even with the
sase-1ex.11 hidden hot spare landed, so the remaining lever is the overlay-docked
prompt bar experiment, recorded as a follow-up on the bead. A cold snapshot never blocks
the first <space>: it opens a blank bar at once and late-prefills only an untouched
session. No budget assertions live in the slow bench by design; the committed regression
gates are the structural tests below, all cheap enough for just check:
- warm
ctrl+p/ctrl+n, first-visit getters, and warm<space>perform zero main-thread MRU reads/writes,list_project_recordscalls,Popens, watcher start/stop, and thread joins (test_launchable_mru.py,test_space_prefill.py,test_prompt_catalog.py,test_prompt_key_perf_smoke.py); - each prompt text-area mount/unmount/worker body runs once, base-first
(
test_prompt_mount_dedup.py); - a stale snapshot generation never overwrites text or reorders an in-flight cycle, and
a cold
<space>followed by typing is never clobbered by a late prefill (test_launchable_mru.py,test_space_prefill.py).
Pager bench¶
The pager benchmark (epic sase-1es) measures open, keystroke, search, memory, leak,
and cold-start cost against a deterministic synthetic corpus over the 500 / 2k / 20k /
100k line ladder. Each case runs in a fresh subprocess with a per-case timeout
(TIMEOUT is reported, never hung), plus a trivial-app floor so CPU numbers read both
raw and above the floor:
just bench-pager
just bench-pager --cases code-sparse,log-dense --max-lines 2000 --no-cold
just bench-pager --cases markdown --ladder 500,2000 --output /tmp/pager-bench.json
The corpus generator is tests/perf/_pager_bench_corpus.py; the harness is
tests/perf/bench_pager.py (slow). test_pager_bench_smoke.py is the fast non-slow
smoke test that runs the smallest case end to end. Baseline numbers are not committed:
shared-host timing is noisy, so record before/after tables in bead notes instead.
The committed tripwires are the regular tests, not the bench: dismissed-view release in
tests/pager/test_view_leak.py, and the strip-cache bound plus the 20k-line
tracemalloc ceiling in tests/pager/test_perf_gates.py.
Targets per phase gate¶
The targets below come from sdd/research/202604/sase_perf_research.md and are restated
here so each phase agent has a single page to check against. A phase is green when the
relevant targets are met without regressing any other span.
j/k highlight p95 < 16 ms
(Stitches next/prev/up10: documented CommitsTimeline carve-out ≤ 30 ms;
unmodified-master baseline was stitches.next 16.47 / stitches.up10 17.95,
conform verification observed stitches.next 20.17 serial / 24.84 under xdist,
agent-pane landing verification observed 26.59 serial / 27.22 under xdist)
key-to-paint p95 < 33 ms
debounced detail paint < 150–250 ms
warm Patch reload, 1k < 100 ms
no-change auto-refresh stall ~0 ms (event-driven path; Phase 7)
large reply first paint immediate plain render, syntax later/optional
Per-phase responsibilities:
- Phase 2 (Patch j/k hot path):
widget.patch_list.update_listcall count drops to zero for j/k navigation;update_highlightp95 < 16 ms at 500 patches. - Phase 3 (data layer): warm Patch reload < 100 ms at 1k specs;
patch.filterp95 should drop materially after the snapshot cache and query context land. - Phase 4 (agent panel + list):
agents.refresh_panel_highlightsandwidget.agent_list.update_highlightp95 < 16 ms at 1k agents. - Phase 5 (incremental loader):
agents.load_from_disknear zero on a no-change auto-refresh. - Phase 6 (artifact + render caching):
widget.prompt_panel.update_display/widget.file_panel.update_displayimmediate first paint on the largest reply fixture. - Phase 7 (event-driven auto-refresh): no-change auto-refresh shows no agents/patch spans firing at all.
Adding a new span¶
from sase.ace.tui.util.trace import tui_trace
with tui_trace("module.name", count=len(items)):
...
Names use dotted lowercase. Counters should be ints / strs only — the emitter falls back
to str(...) for unknown types via default=str, but keeping payloads JSON-friendly
speeds downstream jq slicing.
When a span boundary forces a refactor (most existing hot paths split into foo() →
_foo_impl() so the wrapping context manager doesn't fight indentation rules), keep
both methods next to each other and let the public name stay the trace span name.
Reading the Admin Center Perf view¶
Perf is the eighth view in the Statistics tab of the Admin Center in sase's TUI.
Open Admin Center with #, press 6 for Statistics, then press 0 followed by 8;
[ / ] also cycle to it. The selected Statistics range applies, and g groups
latency by subsystem, provider, or workflow. Perf is global rather than project-scoped:
the project chip stays visible but is marked not applied because the underlying
telemetry and TUI logs do not carry project attribution.
There is no CLI rendering of this dashboard. Use sase telemetry status for store and
configuration state or sase telemetry health for the related traffic-light health
assessment.
What the view shows¶
Five headline tiles summarize Startup, Stalls, Launch, Agent p95, and LLM p95. Startup, stalls, and launch timing come from bounded TUI diagnostic logs; agent and LLM latency come from the local telemetry store. Each tile shows one number, and they are deliberately not the same statistic:
| Tile | Value | Detail line |
|---|---|---|
| Startup | median visible-ready time | sessions in range · slowest initial tab |
| Stalls | stall count | hitch count · worst freeze duration |
| Launch | p95 total launch time | launches · slow stages |
| Agent p95 | p95 agent-run duration | agent runs |
| LLM p95 | p95 LLM invocation duration | error rate · retry rate |
The detailed body contains:
- Startup breakdown — p50, p95, and maximum durations for process→mount, mount→first paint, visible ready, and all surfaces ready, plus the slowest session in the range. The Startup tile itself reports the median visible-ready time and grades it OK below two seconds, warning from two to five seconds, and critical at five seconds or more.
- Stalls & hitches — per-event counts, worst and median duration, recency, suppressed counts, a Freeze records by context ranking, and recoveries. The two tiers are independent watchdogs with different thresholds (hitch at 1.5 seconds, stall at 5, both configurable — see Freeze and hitch capture), so one freeze long enough to trip the stall threshold has already tripped the lower hitch threshold and is recorded as both. The two counts therefore overlap and must not be added together. The tile reports stalls only and names hitches separately in its detail line. Any hitch makes the tile warn; any stall makes it critical.
- Latency & reliability — telemetry-backed p50, p95, maximum, a count, error rate,
and retry rate. The count column is labeled by what it actually counts: LLM
invocations under provider grouping, Agent runs under workflow grouping, and a
plain Count under subsystem grouping. Provider grouping also shows input/output
token and cache data, plus a note that provider rows without LLM invocation samples
fall back to agent runs (which is why those rows read
—for Err%/Retry%). Subsystem grouping omits the Share column entirely, because its rows measure different things and share no denominator; a subsystem with no counter renders—rather than0. The configuredtelemetry.health_thresholdsgrade every row in this panel as well as the Agent p95 and LLM p95 tiles, using the same rules assase telemetry health. - Data & instrumentation — telemetry enablement, selected resolution, store size, raw/rollup counts, write freshness, and one coverage row per diagnostic log. Coverage reports file presence, records in the selected window, earliest retained record, truncation, and unreadable lines. A final line lists the optional probe flags; see Deep profiling and probe flags for how to read it.
The dashboard reads these files from ~/.sase/logs/ by default:
tui_startup.jsonltui_stalls.jsonltui_launch_timing.jsonltui_agent_loads.jsonltui_git_ops.jsonltui_external_tools.jsonl
Missing, disabled, truncated, or partially unreadable sources degrade independently, so the rest of the view remains usable.
The two data sources compute percentiles differently, and the in-app ? help names both
methods:
- Log-derived numbers (startup stages, launch timing, stall medians) use
nearest-rank on the sorted sample, at index
round(q * (n - 1))clamped to[0, n - 1]. - Telemetry-derived numbers (the Latency & reliability rows and the Agent p95 / LLM p95 tiles) are estimated by linear interpolation between cumulative histogram bucket bounds. They are bucket estimates, not exact sample percentiles.
Perf counts come from the telemetry store and the TUI logs, never from the agent-artifact index, so they are not comparable with the run counts on Overview, Projects, or Macros.
Every range except All time also loads the immediately preceding window of the same length; that second load is what the Startup and Agent p95 tiles compare against to show a delta. The other three tiles never show one.
Retention and rollups¶
The view's two data sources age out on entirely different rules:
- Each TUI JSONL diagnostic log is byte-bounded, not time-retained: nothing expires on a
clock, but the oldest records fall off once the file grows. The default limit is 2 MiB
per current file, and rotation preserves exactly one
.1segment before discarding the previous one. SetSASE_TUI_TELEMETRY_MAX_BYTESto override the byte limit. - The telemetry store is time-retained and rolled up: raw samples for 48 hours,
five-minute rollups for 30 days, and hourly rollups for 365 days by default. These
durations are configurable under
telemetry.retention.
So a long lookback reads telemetry at the coarser rollup resolution, and reads only the current plus rotated segment of each JSONL log. All time means everything still retained under those two rules, not an unbounded history.
Deep profiling and probe flags¶
Both probes are off by default and write to their own files under ~/.sase/perf/, which
is a different directory from the ~/.sase/logs/ diagnostic logs the dashboard reads.
The dashboard only names these flags; it never parses what they produce.
SASE_TUI_PERF=1records per-keystrokej/kkey-to-paint samples in~/.sase/perf/tui_jk.jsonlby default. Override the path withSASE_TUI_PERF_PATH.SASE_TUI_TRACE=1records hot-path spans in~/.sase/perf/tui_trace.jsonlby default. Override the path withSASE_TUI_TRACE_PATH; see Trace recorder.SASE_TUI_HEAP=1records periodic top heap allocation sites in~/.sase/perf/tui_heap.jsonlby default. Override the path withSASE_TUI_HEAP_PATH; see Heap sampler.
Each probe records only when its variable is exactly 1. The Probes line in Data
& instrumentation is looser: it prints on whenever the variable is set in the
environment that started the TUI, so a deliberate SASE_TUI_PERF=0 still displays as
on while nothing is being recorded. The same applies to SASE_TUI_TRACE and
SASE_TUI_HEAP. Read that line as "set / unset", and check the value yourself when an
expected probe file stays empty.
sase tui --tmux turns the perf and trace probes on unless the caller has already set
the variable, so SASE_TUI_TRACE=0 sase tui --tmux … (or the SASE_TUI_PERF=0
equivalent) opts out. SASE_TUI_HEAP stays explicit because snapshots add overhead. Use
just view-hints-perf-check for the automated hint-mode regression floor.
Bead history-independence gate (epic sase-1h8, phase sase-1h8.14)¶
The A1 criterion says hot-path bead latency must not move with closed history: on 1x to
8x scaled corpora, p95 of ready, default list, detail read, and a local mutation
moves less than 10%, with warm point read <= 20 ms, active list <= 50 ms, and TUI board
no-change refresh < 100 ms. The gate lives in
tests/perf/bench_bead_scale.py --check-gate (unit-tested by
tests/perf/test_bead_scale_gate.py). list is the paged active-status query that
default sase bead list runs (list_issue_page), not the unbounded list_issues dump;
the TUI op is the cached-snapshot (previous=) path.
# Strict local check: all criteria at <10%, 1x/2x/4x/8x.
just bead-perf-scale -- --check-gate
# CI gate: 1x+4x, runner-noise tolerance, known misses reported not skipped.
just bead-perf-scale-gate
CI (just bead-perf-scale-gate, tolerance 0.5) blocks only on the ready p95 ratio.
The four ratio criteria that still miss, plus all three absolute ceilings, run as
recorded known-misses (--gate-allow): shared-runner wall clocks are
contention-sensitive, and the misses below are real scaling gaps with follow-ups, not
noise. Drop ids off the --gate-allow list as follow-ups land; the strict local run
keeps enforcing everything.
Before/after (medians, ms)¶
Before is the bench-phase baseline (bead sase-1h8.1 notes); after is the
sdd/plans/202610/perf_artifacts/bead_perf_gate.json sweep (runs=5, core pin
5c4033f6, host load ~27, so small-op medians are inflated — clean-window probes in
parentheses).
| op | 1x before | 1x after | 8x before | 8x after |
|---|---|---|---|---|
| detail read | 495 | 37.8 (6.9) | ~4,800 | 60.3 |
| ready | 408 | 34.9 (4.1) | 4,100 | 38.5 |
| default list (paged) | 775† | 65.8 (25.8) | 8,200† | 318.1 |
| note append | 511 | 132.7 (19.4) | — | 257.0 |
| update | 549 | 134.6 (19.1) | — | 258.7 |
| TUI no-change refresh | 33 | 15.6 (9.8) | — | 133.2 |
| remote-backed CLI note | 4,446 | ~1,500 (1,075) | — | — |
audited sase bead read (live store) |
5,000–8,000 | 3,045 | — | — |
† Before measured the unbounded list_issues; after measures the paged default-list
query. Unbounded list still costs ~456 ms at 1x (Python hydration of ~6,900 rows).
Verdict and known misses¶
ratio:ready passes (1.03 at 1x to 8x): ready no longer replays closed history. The
rest miss and stay open with measured breakdowns (each miss is owned by a task bead):
ratio:list(5.1x) andabs:active-list(318 ms at 8x): the paged query touches only active rows, but hydration cost is per row and the active set itself grows 8x with the corpus. Needs bounded serving for unbounded active lists. Owned by sase-1iw.ratio:detail(marginal formally, 3.7x in clean probes) andabs:point-read(25 ms at 4x vs 20 ms): thebead_show_issue_detailbinding itself scales (~6 ms at 1x to ~23.5 ms at 4x) while plainbead_showis flat (0.6 to 0.8 ms), so the relations expansion scans corpus-sized state. Facade hydration adds ~nothing. Owned by sase-1iv. The cause isneighborhood_inin sase-coreread_model/queries.rs, which loads every issue id and scans everylink_provenancerow.ratio:note(1.8x) andratio:update(2.1x; binding-level 2.8x/2.9x in the epic notes): per-mutation cost still grows with stream count. Anstracecount during one facade append (1,383 stat calls total including interpreter startup, under 2,000 streams) shows no full per-mutation stat sweep on that path, so the sweep theory needs revisiting — the follow-up owns the breakdown into admission, publication, and binding overhead. Owned by sase-1iu.abs:tui-nochange(133 ms at 8x, doubling per scale doubling): the no-change path cost is linear in active rows. Needs row virtualization. Owned by sase-1ix.
Contention caveat: this shared host regularly sits at load 20+. One formal sweep was
discarded outright (1x ready p95 88 ms vs 4 ms minutes later on the same store).
Ratios are measured within a single run so both scales share the window, every report
records /proc/loadavg, and borderline verdicts should be re-run before acting on them.