Proof
What has been proved, and what has not.
Here are the 58 things the app must do before we call it finished. Each one is marked as it is tested in the real, installed app; right now, none are. This page is generated from the test results, so it cannot be quietly edited.
Draft: hand-written; the build must generate this from the test matrix. A019 and A025 are summarised from the isolation family; check both against V2.3 before publishing
Why the register tests small steps
Keep AI focused. Keep development in scope.
Given a vague task, an AI still tries to help: it refactors what it was not asked to touch, tries another approach, widens the job. It starts first, looks faster, and is still going hours later, off scope and short of what you asked for. Consensus Logic scopes the work first, then checks after each small step that it is still solving your actual request, and finishes first.
That is the behaviour the scenarios above exist to prove. Until they pass in the installed app, this film is the claim, not the evidence.
Narrated, with captions. Under a minute. It plays only when you press play; nothing on this site moves on its own.
Narrated, under a minute. Captions on screen.
Two kinds of evidence
Tested in the real app, or only in the machinery underneath. We only count the first.
A product pass means the behaviour was observed through the installed application, the thing you would download. A mechanism pass means a test of the machinery underneath passed. Both are real. Only the first counts as proof that the product does what this site says.
| State | Means |
|---|---|
| Verified | Product evidence passed through the installed application. |
| Failed Blocked | A test failed, a dependency is unverified, or a stop is unconfirmed. |
| Pending Unknown | Awaiting approval, or an outcome the controller cannot yet confirm. |
| Not run Not implemented | Honest absence. Never hidden, never counted as a pass. |
The register
The scenarios this site quotes.
Twelve of the 58; the specification calls them acceptance scenarios. These are the rows behind every product claim on the home page and the how-it-works page. Columns for the last real-app run and the build-log entry appear once there is data to show.
| Id | Scenario | Evidence class | Status |
|---|---|---|---|
| A007 | Close the app after a step and resume with artefacts, records, accepted work and spent allowance intact. | product | Not run |
| A008 | The check across a whole document set fails when its sections contradict each other. | product | Not run |
| A011 | Every AI call is reserved against the budget before it starts and settled after; splits and retries never reset the budget. | product | Not run |
| A016 | Both planners write their plans without seeing each other's, from the same evidence; a clean pass needs no manufactured objections. | product | Not run |
| A017 | Stop starts nothing new immediately and shows whether the work actually stopped; if it cannot confirm, nothing else runs. | product | Not run |
| A019 | Project isolation: one project's context is never presented in another project's task. | product | Not run |
| A025 | Project isolation: artefacts and records are stored and listed per project only. | product | Not run |
| A026 | Project isolation is enforced in trusted code, not by a filter on the screen. | mechanism | Not run |
| A031 | Microphone input for describing an outcome. | product | Not implemented |
| A036 | A model's “done” and a green test count cannot unlock the next step; the controller reads the actual output first. | product | Not run |
| A057 | A real disagreement between the planners becomes the user's decision, with the trade-off stated, never a dead end. | product | Not run |
| A058 | Execution cannot be authorised until the plain-English programme summary for that exact plan revision has been shown. | product | Not run |
Scenario wording is paraphrased from the specification for this page. The ids are the specification's own.
Milestones
Each phase of this site opens when a milestone passes its tests. Not before.
The build plan refuses to give calendar dates, because the last two were wrong. This site inherits that refusal. A milestone closes only when the founder accepts the evidence, and only then may the site say what it proves.
| Version | Proves | Status |
|---|---|---|
| M0 | Groundwork: repository, identity, tokens, the harness that every later test runs through. | Closed |
| M1 to M4 | The groundwork in between: saving your work, talking to Codex and Claude, planning, and running steps. | Not run |
| M5 | First working version: documentation tasks run end to end through the installed app. Private beta opens; planning is free; the Founder licence goes on sale. | Not run |
| M6 | Full workflow version: microphone, attachments, both report tabs, all four policy combinations, multi-project. Public beta; Pro and Studio checkout. | Not run |
| M7 | Code tasks version: code and mixed outcomes with release proof. Every scenario verified or honestly marked. | Not run |
Next