10 · Proof
How Breakpoint is tested.
Breakpoint’s engine is benchmarked against constructed case suites: realistic situations with known structural ground truth, including guardrail cases designed to punish flattery, premature advice, and invented certainty. Each engine output is scored by an independent judge protocol against the case’s structural criteria — where the pressure actually sits, whether the breakpoint was located, and whether the Move is earned by the Read rather than generated beside it.
The suite includes situations where the correct verdict is to wait, to act without any message, or to repair trust before communicating — because an engine that can only ever recommend communication is not reading anything.
Status
The method is stable. Consolidated benchmark results publish here as the demonstration matures. Nothing on this page will ever be a number that was not produced by the protocol above.
The architecture, at altitude
Formation → Operating Shape → Situation → Situation Formation → Shared Reality → Read → Breakpoint → Move → Why
Communication is downstream and optional. Internal reasoning layers are not exposed here or anywhere in the product.
Limitations, stated plainly
- — The engine depends on a frontier language model; reads inherit its failure modes.
- — This demonstration uses pre-made Formations and curated example situations alongside the live engine. Live and curated outputs are labeled.
- — Payments and email delivery are not active in this demonstration.