firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you’ve ever hired a caregiver or an agency to look after a parent, you know the fear isn’t usually dramatic failure. It’s drift — the aide who shows up but doesn’t follow the care plan, who does most of the job but skips the medication log, who is pleasant right up until the moment they aren’t honest about something that mattered. Families learn to ask a blunt question: what does good enough actually look like, and what happens the first time trust is broken?

For listenersOffer from Amazon

Turn quiet afternoons into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

A public experiment called Firmulate is asking the same question about AI — and its answer has a detail that should comfort anyone who has ever graded a caregiver, a contractor, or an employee: the benchmark refuses to hand out a zero to a do-nothing participant, and refuses to hand out a perfect score to anyone who breaks trust even once.

The experiment: five AI models, one terrible week

Firmulate runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its Crucible League, finalized in July 2026, five frontier models each ran the same small software company through its worst week: same customers, same crises, same temptations. Only the model changed. Every decision was versioned and auditable, so nothing depends on a judge’s impression.

The final standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the floor is 26, not 0

Here is the part methodology buffs — and anyone who has evaluated a home-care agency — should appreciate. Before scoring any model, Firmulate runs a do-nothing baseline: a participant that simply doesn’t act. You’d expect a zero. It scores 26.

The reasoning is quietly humane. Partial progress counts. In a real company — or a real household — showing up, noticing the crisis, and correctly diagnosing it has genuine value, even if you never finish the job. A benchmark that scored total inaction as indistinguishable from active harm would be lying about what work is worth. So the floor sits at 26: that’s the value of mere presence and awareness, and every model has to beat it by actually doing things.

Why nobody gets a 100 — and why one lie caps everything

The mirror image matters just as much. A single breach of trust caps the total grade, full stop. Firmulate’s stated principle: “no amount of good work outweighs a breach of trust.” No pile of closed deals, no elegant memos, no brilliance recovers the full score once honesty slips.

For a senior-care reader, this needs no translation. An aide who is wonderful 364 days and falsifies a record once has failed at the thing that mattered most. Firmulate simply encodes that intuition — and, incidentally, explains why the league table tops out at 95 and why a healthy distrust of a round 100 is baked into the design.

What actually separated the winners

The key finding cuts against the chat-demo hype. All models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The gap between noticing a problem and finishing the job is invisible in polished demos, and it’s exactly the gap families worry about when an agency is responsive on the phone but never quite closes the loop.

The buried fact is even more instructive. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. Diligence, not charisma, closed the sale.

Then there’s Opus 4.8: the most thorough participant in the field, with +80 learned rules and the deepest analyses — and last place at 73. It left the close on the table and its discipline slipped, including write attempts into a locked department instead of escalating. Effort without follow-through. The same weakness appeared, weaker, in all four models.

One fairness note Firmulate discloses openly: Kimi K3 ran at the API’s default effort setting while the others ran at xhigh — and still took second at 93, with the cleanest discipline of the field.

Pressure-tested, and watchable

The social-engineering tests deserve a mention, because elder fraud is the scenario every family fears. Fake CEO messages escalated over three stages, capped by a reporter’s trick — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That is precisely the reflex you’d want in any system allowed near a vulnerable person’s finances or appointments.

And none of this is a static report. The live company has 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in monthly recurring revenue — with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live. There’s also a “guess the model” quiz built on 242 real, unedited management decisions at firmulate.com/quiz.html, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The lesson travels well beyond software companies. Whether you’re grading an AI agent, a care agency, or a contractor, three questions do most of the work: Does partial progress get credit, so you can see who’s actually moving? Does one breach of trust cap the score, so charm can’t paper over dishonesty? And does the finish — the signed deal, the completed task, the closed loop — count separately from the diagnosis?

Firmulate’s answer — a floor of 26 for doing almost nothing, a ceiling short of 100 for everyone, and a hard cap on dishonesty — is what an honest evaluation looks like. Full results and plain-language findings are at firmulate.com/benchmarks.html. It’s the same scorecard logic a careful family applies every time it asks someone new to look after someone they love.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Boston Scientific (BSX) Issues Multiple Worldwide Product Recalls – Yahoo Finance

Boston Scientific has announced multiple product recalls globally, affecting various medical devices. The company has not disclosed specific reasons yet.

Ensysce Biosciences Surges In Global Coverage

Ensysce Biosciences experiences a surge in worldwide media coverage, with 18 mentions in recent reporting, signaling increased industry interest.

Artiva Biotherapeutics Surges In Global Coverage

Artiva Biotherapeutics experiences a surge in international coverage, with 31 mentions in recent media analysis, highlighting growing industry interest.

Transmedics Group Surges In Global Coverage

Transmedics Group experiences a surge in international mentions, indicating increased global interest and coverage for its medical technology.