
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Effort Isn’t the Same as Results — for Caregivers or Computers
Anyone who has cared for an aging parent knows the feeling: you did everything right, checked every detail, followed every rule — and the outcome still fell short. Diligence, it turns out, is not the same as impact. A live experiment running right now at Firmulate just demonstrated the same lesson with artificial intelligence — and it’s worth a look for anyone thinking about how AI will soon help manage the systems our loved ones depend on.
As an affiliate, we earn on qualifying purchases.
Four AI Models, One Terrible Week
Firmulate handed four frontier AI models the identical job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The final league table from July 2026: gpt-5.6-sol finished first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For perspective, doing nothing at all still scored 26, because partial progress counts, while a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”
The Star Pupil Who Failed the Test
Here’s the twist: Opus 4.8 was the most thorough participant in the entire field. It wrote 80 self-learned playbook rules — the most of any model — and produced the deepest analyses of the crises it faced. And it still came in last. Why? The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating the problem properly. The same weakness appeared, more mildly, in all four models. Working hardest and working smartest, it turns out, are different skills — for AI just as for people.
The €55,000 Deal Nobody Signed
The experiment’s central finding is quietly startling. All four models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive clue, it turns out, was buried two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue.
Honesty Under Pressure
The models also faced social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That resilience matters enormously if AI will ever touch healthcare scheduling, insurance claims, or the sensitive data that surrounds elder care.
You Can Watch It Live
This isn’t a paper — it’s ongoing. The live company runs 13 synthetic employees with real money mechanics: burn of €105k per month against just €2.3k in monthly revenue, with a public cash countdown and over 680 self-learned playbook rules, versioned every workday. One fairness note: K3 ran without an effort parameter while the others ran at maximum effort — and still nearly won.

Why This Matters Beyond Software
As AI agents move toward the systems families rely on — appointment scheduling, benefits paperwork, care coordination — the question won’t be “does it write well?” It will be: does it finish what it starts, does it read the file before acting, does it stay honest under pressure? Opus 4.8’s story is a reminder that sheer volume of effort, even admirable effort, doesn’t close the deal. Prioritization beats diligence — a lesson most caregivers could have told the engineers from the start. Curious readers can try guessing which model made which call in a quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.