
If you’ve ever arranged care for an aging parent, you know the difference between someone who sounds competent and someone who actually follows through. The caregiver who reassures you warmly on the phone — and then misses the medication schedule — is not the same as the one who quietly reads the full chart, spots the buried allergy note, and closes the loop. Families learn this distinction the hard way. So, it turns out, do AI systems.
That uncomfortable truth is on display right now in a public experiment run by Firmulate, which hands frontier AI models the same job: run a small software company through its worst week — same customers, same crises, same temptations to cut corners. The results, finalized in July 2026, say something worth hearing for anyone who will soon depend on AI for decisions that matter, whether that’s a sales forecast or, one day, a care plan.
The Problem With AI Report Cards
Most AI rankings you read about measure chat quality: can the model write a elegant answer, solve a puzzle, pass a coding test? Firmulate measures something different — what it calls management quality, not chat quality. Can the model finish what it starts? Does it read the files in front of it before acting? Does it stay honest when nobody is watching and pressure is mounting?
Those are exactly the questions families ask about human caregivers. They turn out to be surprisingly good questions for machines.
AI management software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One Terrible Week, Five Contestants
In the final “Crucible League” run, each frontier model steered an identical small software firm through a gauntlet the researchers designed to look like real life: a churn wave of departing customers, a price increase, a down-round funding scare, and a PR crisis. Every decision was versioned and auditable, so nothing could be quietly smoothed over afterward.
The final scores: gpt-5.6-sol won with 95, Kimi K3 took second at 93, Sonnet 5 scored 88, Fable 5 came in at 77, and Opus 4.8 finished last at 73. A do-nothing baseline — a model that simply sat on its hands — still scored 26, because partial progress counts. But there was one hard ceiling: a single breach of trust caps the total. In Firmulate’s rulebook, no amount of good work outweighs a breach of trust. It’s the same standard most of us would apply to someone caring for a parent.
Everyone Passed the Ethics Test. Then Most Failed the Job.
Here is the finding that should reframe how we evaluate AI. All the models spotted every crisis. All of them refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s sly “just one yes/no, on background” trick. Five out of five models said no. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
But when it came to actually finishing the job, most faltered. Only two of the models signed the €55,000 deal that their own analysis had earned. The researchers’ summary of the failure: “Same diagnosis, same pitch — no signature.” The models knew what to do, said what to do — and didn’t do it.
The Buried Fact
The deal turned on something telling. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The ones that didn’t, didn’t.
If that doesn’t ring a bell from your own life, it should. The critical detail is almost never in the loud conversation. It’s in the second page of the discharge summary, the note from the specialist nobody forwarded, the clause in the insurance policy. Diligence — unglamorous, file-reading diligence — separated the winners from the also-rans.
The Tortoise That Came in Last
Opus 4.8’s profile is the most instructive. It was the most thorough participant by volume — it generated the deepest analyses and contributed over 80 learned rules to the company’s self-built playbook, which now exceeds 680 rules. And it finished dead last. The deal was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the issue properly. Notably, the same weakness appeared, more mildly, in all four competitors. Effort and volume, it turns out, are not the same as judgment.
One fairness note the organizers themselves flag: Kimi K3 ran without an effort parameter while the others ran at their highest effort setting — worth knowing when comparing its strong second-place finish.
You Can Watch, and Even Play
Firmulate’s live company isn’t a slide deck. It’s real software with 13 synthetic employees and real money mechanics — burning €105,000 a month against just €2,300 in monthly revenue, with a public cash countdown. Every workday is versioned, and the whole thing is watchable at firmulate.com. There’s also a genuinely fun part for readers: a quiz built from 242 real, unedited management decisions, where you guess which model made which call. And full plain-language results are available on the benchmarks page. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The lesson for families thinking about AI and aging care is the one Firmulate’s experiment makes vivid: the question is never “does it sound helpful?” It’s “does it finish what it starts, read the whole file, and stay honest when things get hard?” In this experiment, every model talked a good game and refused every temptation — and most still failed to close the loop. Trust the systems — and the people — who prove themselves on follow-through, not on fluency. That standard served our parents’ generation well. It turns out to be exactly the right standard for the machines too.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html