About
Fifteen years building software. Long enough to know where AI falls over.
Consensus Logic is built by software developers, for software developers, and for the people coming into software development now that AI has made that possible. We use Codex and Claude every day. We have had the sugar high of a build that works first time, and we have had the week that comes after it.
The lived problem
The first build was easy. Keeping control was not.
“It spent hours fixing the problem, said it was fixed, and it was still broken. Then we did it again. I lost days trying to work out what had actually gone wrong.”One of us, on the week that led to this product
We have been building software for over fifteen years, and building with AI for as long as it has been good enough to trust with real work. We are not watching this from the outside. We ship with Claude Code and Codex most days, on Windows, across several products at once, and we use them for the reason everyone else does: they make us faster.
The first builds are astonishing. That is the sugar high, and it is real. The trouble starts later, when changes accumulate, the product stops behaving as expected, and repeated AI-led fixes keep failing to find the real cause.
We have had a build declared done that was not. A large test count passing while the application did not work. Sessions that ran for hours. Quota going into repairs of repairs. Whole weeks gone, with the one person who could have caught any of it busy pasting prompts between three windows, acting as project manager, budget officer and quality control for two AIs.
The tools on offer did not help with that job. A single agent grades its own work. A multi-agent library makes you build the controller yourself. A review tool arrives after the pull request. None of them keeps the authority, the budget and the evidence with the person in charge.
One prompt, twenty hours
The Friday afternoon prompt.
“A shortcut meant to save a few minutes cost weeks of recovery.”One of us, on the afternoon that made the case
There was a working time and mileage tracking system, properly scoped and designed. Late on a Friday afternoon it needed one small change. Instead of writing the careful prompt we normally would, one of us got lazy, turned on the microphone and dictated a quick instruction.
It was going to take a while, so the screens went off and the office emptied for the weekend.
The AI decided to help.
That small request became major modifications nobody had asked for. It ran for more than twenty hours, chewed through a whole ChatGPT allocation, and worked through what should have been several days of development on its own. By Monday the project was broken.
The problem was not capability. It was uncontrolled interpretation. There was no second opinion. No scope check. No checkpoint asking whether the change still matched the intent. No independent review before the next modification went in. The ambiguity in one dictated sentence was read as permission, and nothing in the system was built to stop it.
What that would look like now
Consensus Logic exists to stop one vague instruction becoming twenty hours of unintended development. The work is broken into small steps. Two AIs examine the problem independently. Assumptions are surfaced instead of absorbed. Each change is checked before the next one proceeds, and the budget is set before anything runs. You stay the decision-maker, including on a Friday afternoon.
Why it was built this way
For AI. Against trusting it blind.
We make mistakes. We misread a system we have only half read. We optimise around an assumption that turns out to be wrong. We call something done because the tests went green. Every developer does.
AI makes the same mistakes, faster and with more confidence. That is not an argument against using it. It is an argument for what good engineering teams have always put around fallible people: a second pair of eyes that did not write the work, a limit on how much can go wrong before someone looks, and a record of what was actually checked.
Claude and Codex are extraordinarily capable, and that capability is the whole reason this product exists. The point was never to slow AI down. It was to put a layer of control around it: two independent minds instead of one grading itself, a human decision where the two disagree on something that matters, small checked steps inside a budget you set, and a record that says what was actually proved, including what was not run.
The record matters most. A report a non-developer can read, that never tells a different story from the technical evidence underneath it, is the thing that turns “I think it works” into “here is what changed, where it is, how to try it, and what is still unknown”.
How it is built
The same way it works.
We build Consensus Logic the way it asks you to build: against a specification with fifty-eight acceptance scenarios and a plan that refuses to give dates. One AI creates, the other reviews. Every piece of work closes on a test, and a test that cannot run reports a fail, never a pass. The build log publishes each test's output and the reviewer's verdict, rejections included. The Proof page publishes the register with its real status.
The site follows the same rule. It never claims a capability whose acceptance scenario has not closed, and it names what is not yet built.
The company
Made in Australia.
Consensus Logic is built in Australia. It is licensed for individuals and small studios, and there is no enterprise tier.
The legal entity that owns the mark and the domains will be named here once the trademark clearance is complete. Until then, the honest line is that it is not yet finalised.
Write to hello@consensuslogic.com.
Draft: confirm the fifteen-year claim and the legal entity before publishing
Next