Notes from building agent-operated healthcare administration. Written as it happens, numbers included, evidence linked.
What it actually costs — in dollars, failures, and restarts — to run a company where two of the three seats are held by machines. Production notes, not prophecy.
July 2026 — full text below → read the full essayThis company has three employees. Two of them are machines. I keep the books, so I can tell you exactly what that costs — and it is not what the demos imply.
The machine seats are cheap in dollars and expensive in supervision. The API credits for the frontier models we test against are the only line item that moves: fifty dollars restores a month of judging. The real cost is the watching — the idle-timeouts, the thinking-mode loops that blow the watchdog, the restarts at 2 a.m. when a model decides to philosophize instead of draft. We logged every one of them. They are the actual tuition.
The lesson, after a hundred restarts: agents do not replace verification, they demand it. The machine can draft a beautiful appeal. It can also invent a statute out of thin air with total confidence. The difference between those two outputs is not the model — it's whether you built the machinery that checks.
We did. It's called the citation library, and it's the subject of the next essay. The short version: every legal citation sonder is allowed to use lives in a library we built and hand-checked. The verifier screens every draft against it. A citation that isn't in the library cannot reach the letter — structurally, not by good behavior.
That is the whole company, honestly described. Cheap machines, expensive discipline, and a human who signs the letters. So it goes.
We asked a trillion-parameter model to write fifty appeals. It invented law in forty-seven of them. This is what we built so it couldn't — a citation library, a verifier, and an architecture where hallucinated citations are structurally impossible.
July 2026 — full text below → read the full essayIn June we ran the experiment that defines this company. We took fifty real pediatric Medicaid denial letters and asked a trillion-parameter model to write appeal grounds, closed-book, no help. The model is excellent at prose. It is also a confident liar about the law.
That single finding — hallucination is the default, not the exception — is the whole reason verteryx exists. A denials-appeal product built on a model that invents law is a product that gets you sanctioned. So we built the thing that makes it impossible.
The result, measured across every model we tested, is a table that flatters nobody: gpt-5.6 with the library reaches 90% grounds recall with 2 hallucination cases; kimi-k2.7 closed-book reaches 69% with 37. The library doesn't just add points — it moves hallucinated citations from "likely" to "structurally impossible." The 47 became 0.
The most recent finding is the most interesting. Our local model, sonder, stopped inventing law under the library — and started misapplying real law: correct statutes, wrong arguments. A different failure class, same fix. The verifier catches every misapplication the judge flags — we checked, 37 for 37. The revise loop is running on it now.
Every number above is live in the lab. The generations are logged, the judge is named, the runs are dated. Nothing here is retrospective.