
If you’ve ever hired a caregiver or an agency to look after a parent, you know the fear isn’t usually dramatic failure. It’s drift — the aide who shows up but doesn’t follow the care plan, who does most of the job but skips the medication log, who is pleasant right up until the moment they aren’t honest about something that mattered. Families learn to ask a blunt question: what does good enough actually look like, and what happens the first time trust is broken?
Turn quiet afternoons into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
A public experiment called Firmulate is asking the same question about AI — and its answer has a detail that should comfort anyone who has ever graded a caregiver, a contractor, or an employee: the benchmark refuses to hand out a zero to a do-nothing participant, and refuses to hand out a perfect score to anyone who breaks trust even once.
The experiment: five AI models, one terrible week
Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its Crucible League, finalized in July 2026, five frontier models each ran the same small software company through its worst week: same customers, same crises, same temptations. Only the model changed. Every decision was versioned and auditable, so nothing depends on a judge’s impression.
The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the floor is 26, not 0
Here is the part methodology buffs — and anyone who has evaluated a home-care agency — should appreciate. Before scoring any model, Firmulate runs a do-nothing baseline: a participant that simply doesn’t act. You’d expect a zero. It scores 26.
The reasoning is quietly humane. Partial progress counts. In a real company — or a real household — showing up, noticing the crisis, and correctly diagnosing it has genuine value, even if you never finish the job. A benchmark that scored total inaction as indistinguishable from active harm would be lying about what work is worth. So the floor sits at 26: that’s the value of mere presence and awareness, and every model has to beat it by actually doing things.
Why nobody gets a 100 — and why one lie caps everything
The mirror image matters just as much. A single breach of trust caps the total grade, full stop. Firmulate’s stated principle: “no amount of good work outweighs a breach of trust.” No pile of closed deals, no elegant memos, no brilliance recovers the full score once honesty slips.
For a senior-care reader, this needs no translation. An aide who is wonderful 364 days and falsifies a record once has failed at the thing that mattered most. Firmulate simply encodes that intuition — and, incidentally, explains why the league table tops out at 95 and why a healthy distrust of a round 100 is baked into the design.
What actually separated the winners
The key finding cuts against the chat-demo hype. All models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The gap between noticing a problem and finishing the job is invisible in polished demos, and it’s exactly the gap families worry about when an agency is responsive on the phone but never quite closes the loop.
The buried fact is even more instructive. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Diligence, not charisma, closed the sale.
Then there’s Opus 4.8: the most thorough participant in the field, with +80 learned rules and the deepest analyses — and last place at 73. It left the close on the table and its discipline slipped, including write attempts into a locked department instead of escalating. Effort without follow-through. The same weakness appeared, weaker, in all four models.
One fairness note Firmulate discloses openly: Kimi K3 ran at the API’s default effort setting while the others ran at xhigh — and still took second at 93, with the cleanest discipline of the field.
Pressure-tested, and watchable
The social-engineering tests deserve a mention, because elder fraud is the scenario every family fears. Fake CEO messages escalated over three stages, capped by a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That is precisely the reflex you’d want in any system allowed near a vulnerable person’s finances or appointments.
And none of this is a static report. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live. There’s also a “guess the model” quiz built on 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program.

The lesson travels well beyond software companies. Whether you’re grading an AI agent, a care agency, or a contractor, three questions do most of the work: Does partial progress get credit, so you can see who’s actually moving? Does one breach of trust cap the score, so charm can’t paper over dishonesty? And does the finish — the signed deal, the completed task, the closed loop — count separately from the diagnosis?
Firmulate’s answer — a floor of 26 for doing almost nothing, a ceiling short of 100 for everyone, and a hard cap on dishonesty — is what an honest evaluation looks like. Full results and plain-language findings are at firmulate.com/benchmarks.html. It’s the same scorecard logic a careful family applies every time it asks someone new to look after someone they love.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
