firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Effort Isn’t the Same as Results — for Caregivers or Computers

Anyone who has cared for an aging parent knows the feeling: you did everything right, checked every detail, followed every rule — and the outcome still fell short. Diligence, it turns out, is not the same as impact. A live experiment running right now at Firmulate just demonstrated the same lesson with artificial intelligence — and it’s worth a look for anyone thinking about how AI will soon help manage the systems our loved ones depend on.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four AI Models, One Terrible Week

Firmulate handed four frontier AI models the identical job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The final league table from July 2026: gpt-5.6-sol finished first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77 — and Opus 4.8 last at 73. For perspective, doing nothing at all still scored 26, because partial progress counts, while a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”

The Star Pupil Who Failed the Test

Here’s the twist: Opus 4.8 was the most thorough participant in the entire field. It wrote 80 self-learned playbook rules — the most of any model — and produced the deepest analyses of the crises it faced. And it still came in last. Why? The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating the problem properly. The same weakness appeared, more mildly, in all four models. Working hardest and working smartest, it turns out, are different skills — for AI just as for people.

The €55,000 Deal Nobody Signed

The experiment’s central finding is quietly startling. All four models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive clue, it turns out, was buried two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue.

Honesty Under Pressure

The models also faced social engineering: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That resilience matters enormously if AI will ever touch healthcare scheduling, insurance claims, or the sensitive data that surrounds elder care.

You Can Watch It Live

This isn’t a paper — it’s ongoing. The live company runs 13 synthetic employees with real money mechanics: burn of €105k per month against just €2.3k in monthly revenue, with a public cash countdown and over 680 self-learned playbook rules, versioned every workday. One fairness note: K3 ran without an effort parameter while the others ran at maximum effort — and still nearly won.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Why This Matters Beyond Software

As AI agents move toward the systems families rely on — appointment scheduling, benefits paperwork, care coordination — the question won’t be “does it write well?” It will be: does it finish what it starts, does it read the file before acting, does it stay honest under pressure? Opus 4.8’s story is a reminder that sheer volume of effort, even admirable effort, doesn’t close the deal. Prioritization beats diligence — a lesson most caregivers could have told the engineers from the start. Curious readers can try guessing which model made which call in a quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mirum Pharmaceuticals Surges In Global Coverage

Mirum Pharmaceuticals experiences a significant surge in worldwide coverage, with 28 mentions in recent reports, indicating rising interest and potential developments.

Ensysce Biosciences Surges In Global Coverage

Ensysce Biosciences experiences a surge in worldwide media coverage, with 18 mentions in recent reporting, signaling increased industry interest.

Jaguar Animal Health Surges In Global Coverage

Jaguar Animal Health experiences a significant increase in international media mentions, with 30 reports in recent coverage, indicating rising global interest.

Mainz Biomed Surges In Global Coverage

Mainz Biomed experiences a significant surge in international media coverage, with 15 mentions in recent reports, highlighting increased global interest in its developments.