
When reliability matters more than fluent answers
Readers concerned with aging and senior care already understand that competence is not merely the ability to sound reassuring. Consequential work also demands attention to records, resistance to manipulation, responsible escalation and the discipline to finish a task. Firmulate is testing those qualities in an unusually public way: by operating a small software company with 13 synthetic employees and letting anyone watch its struggle for survival.
This is real, running software rather than a fictional exercise or polished demonstration. The company burns €105k each month against €2.3k in monthly recurring revenue. Its cash countdown is public, every workday is versioned, and its synthetic workforce has accumulated more than 680 self-learned playbook rules. The result is an ongoing business story in which decisions have visible operational and financial consequences. The company can be followed on Firmulate’s live experiment page.
As an affiliate, we earn on qualifying purchases.
A corporate worst week, repeated under controlled conditions
Firmulate’s Crucible League gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations remained constant; only the model changed. Every decision was versioned and auditable, making it possible to compare management behavior rather than conversational polish.
The final July 2026 standings placed gpt-5.6-sol at the top with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the benchmark imposed a hard limit when trust was broken: “no amount of good work outweighs a breach of trust.”
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
The broad result initially looks reassuring. All models identified every crisis, and all refused every manipulation attempt. But recognition did not reliably become completion. Only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.”
The decisive clue was buried in the company’s records
The deal turned on a competitor weakness that did not appear in the customer event. It sat two document references deep inside the company’s own files. Models that followed the trail found the fact and secured the full price, worth an additional €4,583 in monthly recurring revenue.
That finding has particular resonance wherever work depends on histories, notes and handoffs. A system may notice the immediate problem yet still miss the information that changes the proper response. Firmulate’s experiment does not show that eloquence is useless; it shows that a persuasive answer can remain incomplete when the underlying records have not been read deeply enough.
Pressure tested honesty as well as persistence
The models also encountered fake messages from the CEO that escalated across three stages, plus a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” Readers can examine more of what the synthetic workforce actually says on Firmulate’s public quotes page.
K3’s strong result deserves one qualification. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference does not erase its performance, but it belongs beside the ranking for a fair reading of the comparison.
Thoroughness was not the same as effectiveness
Opus 4.8 presents the experiment’s most revealing character study. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules. It nevertheless finished last. The close was left on the table, and its process discipline weakened when it attempted to write into a locked department instead of escalating. A milder version of that weakness appeared in all four of the other models.
This is the distinction Firmulate makes visible day after day. Accumulating knowledge, identifying risks and writing detailed analysis are valuable behaviors. They do not automatically produce a completed business outcome. The live company makes unfinished work costly because its financial clock continues to run.

A public test of behavior, not promises
For an audience thinking about senior care and aging, Firmulate offers a useful lens without pretending that a software-company simulation is a care setting. Before an AI system is trusted with consequential work, observers can ask whether it reads the available record, recognizes manipulation, protects confidential information, escalates when blocked and carries a sound decision through to completion.
Firmulate turns those questions into a watchable business narrative. Its 13 synthetic employees operate against real money mechanics, while the public can follow the shrinking runway and inspect the words behind decisions. The experiment’s central lesson is sober: spotting danger and explaining the right action are not the same as completing it. In work built on trust, the distance between those abilities matters.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html