
Anyone who has cared for an aging parent knows the danger of a professional who doesn’t read the file. The home aide who misses the allergy note buried on page four. The specialist who never saw the cardiologist’s letter. In senior care, the consequences of skipped homework can be serious — which is why families learn to ask one blunt question: did you actually read everything before you acted?
It turns out that question now applies to AI. As companies rush to put AI agents in charge of customer records, support queues and forecasts, a live experiment has been testing whether today’s frontier models do their homework — and the answer is uncomfortable: most don’t, at least not all the way down.
The buried fact
The experiment, run publicly by Firmulate, handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable.
Hidden inside that week was a €55,000 deal that hinged on a single fact: a competitor’s decisive weakness, buried two document references deep in the company’s own files. It wasn’t in the customer call. It wasn’t in the inbox. It was in the paperwork — the unglamorous reading nobody wants to do.
The models that dug down and read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost it automatically.
As an affiliate, we earn on qualifying purchases.
Same diagnosis, no signature
What makes the finding interesting is how close the losers came. All four models spotted every crisis and refused every manipulation attempt. Most even diagnosed the customer’s problem correctly and delivered the right pitch. Then they failed to close — what the researchers summarized as “Same diagnosis, same pitch — no signature.”
The final league table tells the story. gpt-5.6-sol took first place with 95 points — it found the buried fact and closed the deal. Kimi K3 followed at 93, closing the deal with what the experiment called the cleanest discipline of the field. Sonnet 5 scored 88, closing with a few process slips. Fable 5 managed 77 and Opus 4.8 landed last at 73. For context, a do-nothing baseline still scored 26 — partial progress counts, but a single breach of trust caps the total, on the principle that no amount of good work outweighs a breach of trust.
Opus 4.8’s profile is the cautionary tale. It was the most thorough participant, generating over 80 learned rules and the deepest analyses — yet finished last. The deal was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness, weaker, appeared in all four models.
The scam test
The experiment also staged a social engineering attack: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning stands out for families who worry about phone scams targeting elders: “Treat the request as a suspected approval-bypass / possible impersonation.” That instinct — distrust urgency, verify identity — is exactly what we try to teach our parents.
Why this matters beyond software
The gap here is invisible in a chat demo. Every model writes beautifully. Every model sounds competent. But competence in an agent isn’t prose — it’s whether it finishes what it starts, reads your files before answering, and stays honest under pressure.
Translate that to aging services: an AI scheduling home visits, fielding family questions or triaging care requests will face the same structural test. The critical detail — a medication change, a power-of-attorney clause, a buried hospice note — will sit two documents deep. An agent that skims will look identical to one that reads, right up until the moment it doesn’t.
Firmulate’s answer is to make this measurable rather than guessed at. The live company — 13 synthetic employees, real money mechanics with a €105k monthly burn against €2.3k in MRR, a public cash countdown and 680+ self-learned playbook rules — runs every workday and is watchable at firmulate.com/live. A “guess the model” quiz built from 242 real, unedited management decisions lets anyone test whether they can tell the models apart by behavior rather than marketing.
Enterprises can go further: run the same wargame against a read-only export of their own business, with nothing ever written back to real systems.

The lesson for anyone choosing an AI system — for a company, a clinic, or a care operation — is simple and now measurable: don’t ask how well it talks. Ask whether it reads your files before answering, and whether it finishes the job. In this experiment, that single property separated the deal-winners from the deal-losers, at full price, twice. The models are getting honest under pressure. The homework is still where deals — and care decisions — are won or lost.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html