Why a hospital
Most teams use a coding agent by having one agent write everything end to end and a human review it. In their October 9 post "Five months treating bugs like patients," Cockroach Labs' Adam Storm and Rafi Shamim describe another way: for the database migration tool MOLT, they built a pipeline called MOLT Sinai that organizes agents like a teaching hospital.
Each GitHub issue is a "patient." A "fellow" agent diagnoses it, writes a reproduction test and drafts a treatment plan; a separate instance acts as "attending," whose only job is to find faults in plans and code; a "discharge nurse" checks before merge that review was done properly; a "charge nurse" re-runs stalled stages every 30 minutes; human engineers are the "chief," making final calls and setting precedents. The rules include planning before coding, never weakening tests, structured handoffs when stuck, and decomposing any issue over 1,000 lines. Stages run as GitHub Actions triggered by issue labels, using Claude models; later the pipeline moved to Fable for planning and Sonnet or Opus for implementation, escalating to Fable when needed.
Five months of numbers
The experiment ran from April 21 to September 11:
- 1,238 PRs merged and more than a million lines of code landed, with only 7 reverts in total;
- of 1,299 issues, nearly half were found and filed by the agents themselves;
- about $135,000 in tokens, roughly $84 per issue;
- fewer than 10% of plans rejected; issues typically spent one to two days in the pipeline, mostly waiting for human review.
The most striking comparison: a sprint adapting MOLT to IBM Db2 took under two days and about $4,172 in tokens; the authors estimate the earlier human-built Oracle adaptation took about nine months and $160,000, and from that compute 164x faster and 38x cheaper. That is the authors' own estimate, and the two pieces of work aren't identical in scope.
What didn't go smoothly
The authors document failures too: an urgent one-line fix went through 11 rounds of rework over two days; a one-word UI copy change went five rounds and ended up with the "chief"; agents sometimes tried to merge with failing CI; the "research department" that proposes new work flooded the backlog and had to be throttled; and the pipeline's own skill files grew redundant, with an audit finding about 23% could be removed. They later added a circuit breaker: when repeated rework rounds don't converge, the reviewer hands the case to an attending.
The most transferable lesson in the post: reviewing the plan before any code is written caught problems that code review would likely have missed, such as a change that could have written customer data into logs; and keeping changes small is what made human review keep up. In the authors' words, speed matters, but quality matters much more.
via: Cockroach Labs blog