No definition of done
The agent optimises for "the tests pass" or "the user stopped complaining". Nobody wrote down the real outcome, so nobody can check against it.
The Foundry turns your business — mapped as a Business Genome — into working software, and proves every change with the Eval Harness before it goes anywhere near production.
Every team can now produce code quickly. What most can't do is say why a change was made, whether it's correct, or what it costs. Four gaps cause almost all of it.
The agent optimises for "the tests pass" or "the user stopped complaining". Nobody wrote down the real outcome, so nobody can check against it.
Every change is treated as equally risky. So either everything is waved through, or everything queues behind one senior person.
Each person's AI rediscovers the business from scratch. The tenth internal tool doesn't know the first nine exist.
Sessions run and are thrown away. The same failure is fixed forty times, in forty different ways, and nothing gets better.
Work enters as a defined signal and leaves as a defined outcome. People appear where judgement is needed: setting the outcome, and looking at what's risky.
Your business mapped: processes, systems, decisions, KPIs.
→ shared contextThe outcome written down, with the tests that will prove it.
→ definition of doneBuilt from the component library first, new code only where needed.
→ working changeThe Eval Harness runs stubs and scenarios against the spec.
→ evidenceRisk class decides what a human reviews and what flows through.
→ decisionReleased with owner, approvals, spend caps, audit trail and rollback.
→ live outcomeEvery failure becomes a test. Every pattern becomes a component.
→ better next timeThe Eval Harness is how the Foundry knows a change is right. It ships with every production line and stays with you afterwards, as your own regression suite.
A stub stands in for anything the code depends on — an ERP, a bank portal, a model, a colleague's service — and answers exactly as the real thing would, including on its bad days.
A scenario runs a whole business journey end to end, the way your team would describe it, and checks the outcome the business actually cares about.
How hard the gate is depends on the risk class of the change. Low-risk work flows through on evidence alone; anything that touches money, customers or compliance waits for a person.
| Check | Question it answers | Evidence |
|---|---|---|
| Functional | Does it do the job described in the spec? | scenario pass rate |
| Accuracy | Is the output right against an agreed benchmark? | scored golden set |
| Integration | Does it call the right systems, the right way? | stub contract tests |
| Security | Can it reach only what it's allowed to? | permission + injection tests |
| Safety | How does it behave on ambiguous or hostile input? | red-team suite |
| Regression | Did this change break anything that worked? | full suite on every change |
| Performance | Is it fast enough at real volume? | load run |
| Cost | What does one transaction cost to run? | cost per run vs budget |
| Business impact | Did the number we agreed actually move? | before and after |
Most teams start every project from a blank page. The Foundry starts from what we already know about your business, and gives back more than it takes.
Nine functions across your real end-to-end processes, people, systems and numbers.
That understanding becomes a structured, scoreable map of the company that software can be built against.
Specs, components and tests all reference the Genome, so nothing is invented twice.
Running software feeds real data back: what broke, what cost too much, what changed.
Each task runs on the model and platform that suit it. Swap providers without rebuilding business logic — the tests prove the swap is safe.
Specs, components, stubs, scenarios, gates and routing rules are all versioned. "Improvement" is measurable, not a feeling.
You own the code, the Genome, the Eval Harness and the data. Your team can run and change all of it without us.
A short Business MRI to find where the work actually hurts, and agree the one outcome worth proving first.
One loop, built and running: spec, components, Eval Harness, gate, release. Part of our fee rides on the result.
Your own lines, your component library, your test suites, your gates — with your team trained to run them.
Most clients ask us to keep operating it. Regulated ones take it in-house. Either way, the Foundry stays theirs.
They turn a ticket into a pull request. They don't know your processes, and they can't tell you whether the outcome was right for your business.
Slides, target operating models and a plan. Then the building, testing and running is someone else's problem.
Method, gates, stubs and reuse take a year to get right. We bring them on day one and hand them over.
Thirty days, one outcome, evidence you can check. If it doesn't move the number we agreed, part of our fee goes unpaid.