Measure the system.
Don’t confuse the measurements.
Runtime health, operational activity, continuity quality, and experiments answer different questions. They stay separated here so a useful number never gets mistaken for a universal score.
Is the machinery behaving correctly?
What durable work exists right now?
A privacy-safe window through the sudofx boundary.
Shows lifecycle evidence only: which application used sudofx, when it happened, which boundary stages completed, and bounded context size. Prompt, response, research content, and payload data are never shown here.
Waiting for application invocation evidence…
The next governed application invocation will appear here without exposing its contents.
Which applications are using the engine, and how?
Derived from generic governed application actions and invocation lifecycle evidence. This view is disposable; SQLite remains authoritative.
Waiting for durable application evidence…
No application-specific meaning is inferred here.
Did a fresh intelligence receive enough faithful context?
These are measurements about one continuity experiment, not general model-quality scores.
Bounded context versus full source context
Current durable observations created before this metrics contract may not contain byte aggregates. New cycles preserve only the safe aggregate measurements.
Where is the continuity stress test?
Application meaning stays local.
WAKE✳︎ owns research-specific metrics such as cycles, evidence collection, notebook progress, and research funnels. sudofx exposes only generic authority/runtime measurements here.
Open WAKE✳︎ metrics →Evidence is not authority.
A fast replay, a passing semantic review, or a high compression ratio is evidence about a specific boundary. None of those measurements gets to rewrite the durable record.