AI-agent orchestration · free kit (MIT) · hands-on help
Put AI agents to work and keep the last word.
Guildwork is a discipline built one incident at a time, and upgraded by every new one: a contract before every task handed to an AI, a witness while it runs, a closeout after — and a launcher that refuses to start what isn’t in order. Running today on real software (seven agents, twenty-five merges a day), it holds a law firm, an agency or a support desk just as well.
The kit is public and free. Installing it where you work is my job.
Interactive demo · nothing to install
Spend a day in charge
Five jobs, agents that do them, and you. One mistake stopped before it went out, one decision that is yours, and a tally at the end: what got delivered, and how many times you clicked. Two minutes, with code or without.
Start the day →Measured on 9 September 2026 · every figure on this page is real
Two doors, one model — pick yours
You have code
A team, one repository, agents crowding into it
Claude Code, Codex CLI, Grok or others, on code that matters. The launcher that refuses, the witness, the closeout, the board — and thirty-eight days of numbers, still counting.
You have no code
A company, its processes, and AIs already at work
A law firm, an agency, a restaurant, a support desk. What you could already hand to an AI, what it would give you back in hours, and the first process under contract within a week.
Four failures no prompt cures
Several AI coding agents — different vendors, different harnesses — working on one repository at once produce the same four kinds of damage, in every team that tries it.
01 · THE SECOND WRITER
Two agents write the same file
and the human becomes the merge tool — the one job the agents were supposed to remove.
02 · THE SILENT SUBSTITUTION
An agent runs under undeclared parameters
another model, another effort level, another prompt, and nothing records the substitution.
03 · THE LOST SESSION
What an agent learned dies with its context
so the next session re-derives it — or works from a stale summary, with confidence.
04 · THE MESSAGE BUS
The human carries instructions by hand
from one agent to the next, all day, which is exactly what the agents were bought to stop.
Read it for where it stops — and how it moves
A diagram of agents usually shows who calls whom. This one shows where a mission can be refused, and what carries it: ten checks before a session starts — seven stable contract conditions, three guards hardened in operation — a witness and a human gate before anything lands, three exit codes decided by durability alone, a register that sends every incident back into the next contract, and a watcher drawn twice: what it does today, and the receive–deliver–wake loop that still has to earn activation.
Drawn on Lantern, the fictional project the kit is demonstrated on — the mechanism, the checks and the exit codes are the real ones. Open the interactive version and hover a block to isolate its path.
What a refusal looks like
Watch one arrive → — inside a working day, in two minutes.
The launcher read contract #118, fetched the governance at the commit the mission would run on, and stopped before creating anything. Three days earlier an attempt at the same lot had been refused at review; its branch survived on the remote because nobody wrote a disposition for it. Silence is debris — and the launcher saw it before a second writer existed.
C:\repos\lantern\.venv\Scripts\python.exe C:\repos\lantern\tools\mission_launcher.py 118 --dry-run Mission #118 — #118 — [MISSION] Dashboard tiles keep their last reading while the device bridge reconnects https://github.com/lantern-workshop/lantern/issues/118 primary checkout: C:\repos\lantern governance policy: 3f9c2d1a7b8e4c05d6f1a2b3c4d5e6f708192a3b (origin/develop, fetched at launch) run: launch — creates the branch and the worktree the contract declares REFUSED — 1 refusal(s). Nothing was started. condition 4: a second writer would be created branch 'feature/dashboard-stale-reading' already exists on origin -- a writer already holds the branch this Mode 'one writer' mission declares, and creating it again locally with 'git worktree add -b' would put a second, diverging line on the same name; under 'one writer' that is the two-writer failure itself
The whole transcript — the refusal, the disposal, the run that passed, and what --list showed forty minutes in — is in the kit. The doubled number in the banner is real, known, and cosmetic: the title carries its own number by convention.
The board, read for what it owes you
Twenty-five merges a day are not readable from an issue list. The Product Owner opens one page instead — the page shipped in the kit under templates/board/ — and it answers three questions, in this order, and refuses to guess the answer to any of them.
Owed to you
What only you can decide — and has actually arrived
A reserved gate is not a wait. gate:po says who will decide; it does not say the decision is on your desk. The board separates the two, and lists only the contracts whose delivery is open and carries the line Mission: #N delivers.
If a contract is not on this list, nothing is waiting on you. That is the sentence the page exists to make true.
Moving without you
What the architect merges on the witness's word
Deliveries whose gate is delegated pass without your signature. They are shown anyway — a delegation you cannot see is not a delegation, it is a blind spot. You watch them go by; you have nothing to do about them.
Visible and actionable are not the same thing, and the board is explicit about which is which.
Refused to guess
Whatever fits no rule, counted on the face of the board
A pull request with no Mission: line. A delivery naming a mission that is already closed. A lot carrying no chantier: label. A dashboard that quietly files these away is telling you a story; this one puts them in a counter you cannot miss.
Those counters are the measure of what the governance is not yet doing. They are meant to read zero.
The board as it runs, on the fictional Lantern example — open the demonstration. It is frozen and says so on its face: the controls are shown, not live. What the kit publishes today is the board reader, under templates/board/; the host reading and the decision path are newer, and reach the kit when they have stopped moving.
For someone who does not write the code, this is the difference between “we use AI” and “we know what the AI did.” Not a productivity chart — a list of what is owed, to whom, and what nothing yet accounts for.
What it runs on
A Windows desktop application in Python with a 13 155-test suite, measured on 8 September 2026. One human — the Product Owner — who holds the hardware, the product verdict and the merge to the release branch. Every line of code, test and governance document since early August 2026 has been written by AI seats under these contracts — and still is.
The sober reading, because it is the one that makes the first credible. The same period produced forty-seven incidents serious enough to earn a rule, most of them inside one week. Most seats the launcher can start are unknown on the hardware property the queue would most like filled. Neither command-line harness enforces read-only from its own flags, and the governance says so rather than pretending. A system that can say I do not know in the places it does not know is the only kind whose yes is worth anything.
Five weeks, one repository, one human
The data, as a table
| day | merges | entries |
|---|---|---|
| 01/08 | 11 | — |
| 02/08 | 20 | — |
| 03/08 | 24 | — |
| 04/08 | 17 | — |
| 05/08 | 20 | — |
| 06/08 | 13 | — |
| 07/08 | 14 | — |
| 08/08 | 8 | — |
| 09/08 | 5 | — |
| 10/08 | 2 | — |
| 11/08 | 6 | — |
| 12/08 | 1 | — |
| 13/08 | 7 | — |
| 14/08 | 2 | — |
| 15/08 | 2 | — |
| 16/08 | 6 | — |
| 17/08 | 16 | — |
| 18/08 | 4 | — |
| 19/08 | 9 | — |
| 20/08 | 12 | — |
| 21/08 | 15 | — |
| 22/08 | 7 | — |
| 23/08 | 2 | — |
| 24/08 | 4 | — |
| 25/08 | 14 | — |
| 26/08 | 14 | — |
| 27/08 | 13 | 10 |
| 28/08 | 34 | 29 |
| 29/08 | 21 | 36 |
| 30/08 | 31 | 1 |
| 31/08 | 32 | 13 |
| 01/09 | 19 | 5 |
| 02/09 | 19 | 1 |
| 03/09 | 4 | 5 |
| 04/09 | 34 | 17 |
| 05/09 | 41 | 6 |
| 06/09 | 38 | 12 |
| 07/09 | 20 | 13 |
The tempo. Before the contracts, 241 merges in twenty-five days — under ten a day, with the human carrying instructions between seats. After them, 320 in thirteen — twenty-five a day. The seats changed at the same time as the governance, so the doubling is not the governance's alone. What is the governance's is that one person held that tempo without becoming the merge tool.
The cost of learning. Every register entry is a rule a finding paid for. In the first week of the contracts the team wrote one for every two merges; in the second, one for every three, at the same pace of merges. Before 26 August the findings were not counted at all — the first thing the system did was make them visible.
Fifteen documents, numbered in reading order
Start with 00 for the shape and 12 for why it is shaped that way. If those two earn a third, read 05.
00The operating model
01The mission contract
02The delivery contract
03The label taxonomy
04Capabilities and routing
05The launcher
06Session entry and exit
07The closeout tool
08Continuity
09The findings register
10Effort and execution parameters
11Changelog fragments
12Incidents — and the rule each one paid for
13By the numbers
14The adoption path
templates/ · examples/The drop-in files, and one fictional mission followed end to end
Where this is going
The kit is the part that has stopped moving. Three things are still moving on the project it comes from, and each will reach the kit the way everything here did — once a real incident has paid for it.
The board becomes the control. Not a report on the agents: the page where the one decision that is yours is taken, and recorded by the same click. Twenty-five merges a day are not survivable any other way.
A watcher that reads now, and routes next. Today it observes authorised events and fresh host readings every two minutes without starting a model. After its adapter canaries, it can receive the route, deliver the canonical contract and wake the named seat — without inheriting any authority.
The same seven states on every board. Running, holding the screen, next, for you, landed, broken, expired — whatever the board follows. The names adapt to a project; the positions and the hues do not.
What adopting it looks like with help
The three tools are specified precisely enough to audit or re-implement, and they are not shipped. Installing this on a repository, adapting the vocabulary to a team's own seats and workstreams, measuring the capability table on their machines, writing the tools against their harnesses, and running the first two weeks of missions alongside them — that is the work, and it is the work I do.
Start here
Governance audit
Two days. I read the repository, the board and one week of agent sessions, and write down what will break before it does — then what you could run on top, once it stops breaking.
- the capability table for your seats, measured, not guessed
- the incidents your setup is currently exposed to, named
- a dry run of the ten current launch checks over your existing board
- the three fixes to make before all the others, ordered by damage avoided
- what the install would gain you, costed — and what it will not fix
- a written report your team keeps, whatever happens next
€2 500 · fixed priceDeducted from the installation if you go on.
Book 30 minutes→The install
Governance installed
Two weeks. The contracts, the vocabulary and the three tools, written against your harnesses and your board — then the first missions run with your team, not for them, until the rules are theirs.
- issue form, delivery template, label taxonomy, role profiles
- launcher, session cycle and closeout, written against your CLIs — the code is yours
- the capability table, measured on your machines
- the workstream board, wired to your repository
- two weeks of real missions run with your engineers — the training is the work itself, not a slide deck
- the findings register opened on day one, ruled on with you until the last
€12 000 · fixed priceFor one repository and up to four seats; beyond that, the audit sets the figure. Remote, or on site in Île-de-France.
Book 30 minutes→After
Standing architect seat
Three days a month. The register gets ruled on, the rules follow your incidents, and the governance keeps up with the harnesses as they change — you do not inherit a frozen system the day I leave.
- the findings register reviewed at every board sweep
- contracts and routing kept current as seats change
- the tools kept up when a CLI ships a release that breaks an assumption
- your new joiners walked through the contracts already in place
- a monthly written reading of what the numbers say
€3 500 / monthThree months minimum, then monthly. Extra day: €1 200.
Book 30 minutes→The thirty minutes are free and commit you to nothing — we look at your setup, and you leave with an opinion even if you go no further.
The repository alone
- fifteen documents, the issue form, the pull request template, the label recipe, eight role profiles, the journal format
- the specification of the launcher, the session cycle and the closeout tool — precise enough to audit, and not shipped
- the incidents, anonymised, with the rule each one paid for
- a fictional mission followed end to end
- what it costs: the three tools — seven thousand five hundred lines on the project the kit comes from — written against your harnesses, and the incidents you will meet on the way
With me
- the three tools written against your CLIs and your board, by the person who wrote the originals
- your seats' capability table measured on your machines, not copied from mine
- the vocabulary adapted to your workstreams before the first contract is written
- the first two weeks of missions run with your engineers — the judgment of which rule applies to your harness is what the incidents cannot transfer
- a register that starts on day one, with someone who has already ruled on more than a hundred and fifty entries
Three questions before saying you do not need this
Two things are true at once: your teams could hand an AI far more than they think, and nobody at your company keeps a record of what they hand it already. The second is what stops the first — you do not accelerate on ground you cannot see.
I am not trying to convince you. These three questions are enough, and you already know the answers.
Who, in your company, had an AI do something this week — and where is it written down?
A quote, a customer reply, a meeting summary, a contract clause. If the answer is "in their ChatGPT history", it is written at the vendor's, not at yours.
If that person left tomorrow, what would remain of the way they do it?
The employee who "knows how to handle" the AI has a method. It lives in their head and in their prompts. It leaves with them, like the Excel macros of old.
The last time an AI got something wrong at your company, how did you find out?
From a customer. By chance. Three weeks later. Or — the most frequent answer, and the most honest — you did not.
If the answer to the third is "we did not", you already have the incident. What you are missing is the register — the place where every error becomes a rule instead of happening again.
What it earns you, and what it spares you
The first two are what it earns. The other four are what it costs when nobody holds the thread. In none of them did anyone cheat or cut a corner. Under each, the rule that makes the difference, in one sentence.
The firm
Two contradictory answers to the same client, one day apart
Two associates, two AI assistants, two ways of asking the question. On Monday the client gets one position. On Tuesday he gets the opposite, on the same matter. He calls the partner, who finds out at the same moment he does.
Nobody cut a corner, nobody went around anything. What is missing is not one more rule: it is knowing, for every answer, what was used to write it and who read it back. Once that is in place, both of them can put their AI assistants to work at full speed.
What was missing, in three words
The written request. What is being asked, from which documents, what must come out. In one place for both, before anyone starts.
The sign-off. Someone other than the author reads it before it goes, a colleague or a second AI assistant, and their name stays on it.
The register. The first contradiction, written down in one line, would have become the rule that prevents the second.
The hotel
Group enquiries go out the same day
Ten or so a week, priced between two emergencies: the quote went out three days later, and the client had booked elsewhere. An AI assistant prices them all on this season's rates; the manager reads and signs before noon.
Every quote says which rate card was used, and nobody but the manager sends it.
The online shop
Two thousand products nobody could find
Two thousand products online with no description: invisible to anyone searching for them, and it has been that way three years. An AI assistant writes them from the manufacturer's sheets, in batches of fifty, and a person approves each batch.
Every page says where its claims came from, and a person says yes before it goes live.
Human resources
A candidate's file left the company
A colleague pastes the CV and her interview notes into the tool she uses at home to write the rejection. Salary expectations, personal impressions. If that candidate asks what was kept about them, nobody can answer.
Everyone knows which tools are allowed, and what leaves the company leaves a record.
The support desk
A refund promised outside the rules
The AI assistant answering messages offered a gesture nobody had authorised. The client accepted it, so it is owed. Found three weeks later by accounting, and nobody can say which rule it thought it was following.
The AI assistant knows what it may promise, refuses the rest, and says that it refuses.
The agency
An invented figure, published under the client's name
The figure was credible, the source did not exist, and the text stayed online for four days under the client's name. The client is the one who spotted it. Nobody at the agency could say what the AI assistant had been asked, or by whom.
Nothing is published until a person has read it, and what was used to write it is kept.
Management
The method went on leave for three weeks
The executive assistant has the board minutes produced by an AI, her way, and the result is good. She is away. Nobody knows what she asked for or from which documents, and the next minutes have to be redone.
Every piece of work leaves a record the next one can read. The method stays in the company.
One page, and you know where everything stands
It is one page. It reads what your teams had an AI do, a quote, a reply to a client, meeting notes, a contract, whatever tool produced it, and it sorts everything into three piles. What is in them changes from one trade to the next. The three piles do not.
Owed to you
Three things today, not thirty
The quote above your threshold. The reply to an unhappy client. The contract an AI assistant drafted. Each carries the name of the person who has to say yes, and reaches them only when it is truly ready.
If it is not on this list, it is not waiting on you. You are no longer the person everything has to go through.
Moving without you
Done, without you, and you can read it
Routine answers, confirmations, the summaries nobody needs to argue about. They go out without your signature, and they stay on show.
You do not approve them. You keep the list of everything that went out in your name.
Refused to guess
What nobody is allowed to guess
A document nobody can trace. A piece of work nobody owns. A reply to a request that was already settled. These are not filed away quietly: they are counted, and everyone sees them.
That counter says where your organisation is not holding yet. It is meant to read zero.
A fictional firm, its quotes, its client replies and its minutes: open the demonstration. It is frozen and says so on its face, the buttons are shown and not live. There is nothing to translate: these are your own objects, and the decision is taken on the page rather than somewhere else.
This matters more without code, not less. A development team at least has a version history: someone can reconstruct who changed what. A firm whose AI assistants draft, answer and summarise all day has no history at all — until someone gives it one.
Where this is going
One page that reads what your people hand to an AI — quotes this year, minutes and client replies the next — and sorts it into the same three piles, under the same states from one trade to the next. The person who signs sits in the consultant's chair, not the executor's: the AI brings the finished thing and the one decision that is theirs, and everything else leaves in plain sight.
Offers and prices
Most companies I meet badly underestimate what they could hand over — and overestimate what they would have to buy to do it. So we start by looking at your real work: what can be handed over right now, what that gives you back in hours, and what must stay in human hands. Nothing to buy, and the document stays with you whatever happens next.
Start here
Scoping day
One day on your premises, looking at how you actually work — not only at what you already do with AI. I leave with what can be handed over right now, what that gives you back in hours, and what must stay in human hands.
- the tasks that can be handed to an AI today, counted in hours a week
- the ones that cannot yet — and what is missing for them to be
- the map of your real AI uses, including the ones nobody declared
- the first contract written, on the process worth handing over most
- a document your team keeps, whatever happens next
€1 200 excl. VAT · fixed priceDeducted from the next step if you go on.
Book 30 minutes→The install
First agent under contract
One week. The process chosen for what it costs you, the agent that takes it over, and the three guardrails that make that safe — built with the team that will run it, not in their place.
- the agent itself: what it reads, what it produces, with which tools — built in front of you, not a black box
- the contract: what it may do, what it must refuse, what it must name
- the sign-off: who reads it back, and what never leaves without one
- the register: every error becomes a rule, written, the same day
- five days of real cases run with your staff — the training is the work itself
- the process measured before and after: what it cost in hours, what it costs now
€4 500 excl. VAT · fixed priceFor one process and one team; the second is priced at the scoping day. On site in Île-de-France, or remote.
Book 30 minutes→After
One day a month
One day a month, on site or remote. The register gets ruled on, the contracts follow your incidents, and the perimeter widens one process at a time.
- the register reviewed, every new rule written with you
- contracts adjusted when tools, people or your prices change
- the next process put under contract as soon as the previous one holds
- what your agents actually cost, tracked month by month: subscriptions, tokens, review time
- your new hires walked through the contracts already in place
- a written reading, every month, of what the agents did — and refused
€1 200 excl. VAT / monthThree months minimum, then monthly. Extra day: €1 200 excl. VAT.
Book 30 minutes→The thirty minutes are free and commit you to nothing — you describe your week, I tell you what I see in it, even if you go no further.
Where the discipline comes from. It was built, incident by incident, to hold a seven-agent AI team on real software with a single human in charge — twenty-five deliveries merged a day, every error turned into a rule, all of it public and measured. It is Guildwork. At your company there will be no repository and no code; there will be the same written requests, the same sign-offs, the same register.
152rules written in five weeks, each paid for by a real error