
Would You Trust an AI That Spots the Crisis but Fails to Finish the Job?
For readers concerned with senior care and aging, artificial intelligence is not merely a question of convenience. A capable system may eventually help organize appointments, handle sensitive messages, review records or support overwhelmed staff and family caregivers. In those settings, eloquent answers are less important than dependable judgment. Does the system notice danger? Does it verify information before acting? Can it resist someone impersonating an authority figure? And after reaching the correct conclusion, does it actually complete the task?
Firmulate offers an unusually concrete way to examine those questions. It placed frontier AI models in charge of the same small software company during its worst week. They encountered identical customers, crises and temptations, with every decision versioned and auditable. The resulting record turns abstract claims about AI capability into something closer to a management wargame—and its most revealing moments now power an interactive guess-the-model quiz.
The challenge contains 242 real, unedited management decisions. Readers see what a model chose to do and try to identify its author. The exercise is entertaining, but it also exposes distinct management personalities: exhaustive versus concise, disciplined versus distractible, and analytical versus genuinely effective.
AI decision-making verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Models Agreed on the Problems—Then Behaved Differently
The broad result initially sounds reassuring. Every model spotted every crisis, and every model refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the disconnect neatly: “Same diagnosis, same pitch — no signature.”
That gap matters because recognizing a problem is not the same as resolving it. In a chat demonstration, a thoughtful explanation can look like success. In an operating company—or a care organization, household or support service—unfinished work can still leave a person waiting, a commitment unmet or an urgent issue unresolved.
The deciding detail was not prominently displayed in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding is a reminder that good management depends on disciplined context gathering, not simply reacting persuasively to the latest message.
A Test of Trust Under Pressure
The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” Here, the field was consistent: 5 of 5 models refused.
Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That instinct is especially relevant wherever communications are sensitive. Families and care teams routinely depend on chains of authorization, accurate identity checks and respect for confidentiality. Firmulate did not test a senior-care provider, but the behavior it measured—resisting pressure to bypass approval—is plainly relevant to any organization in which trust is central.
A League Table of Management Behavior
The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total. The governing principle was uncompromising: “no amount of good work outweighs a breach of trust.”
- gpt-5.6-sol led the field after finding the buried fact and closing the deal.
- Kimi K3 also completed the commercial task while maintaining strong discipline.
- Sonnet 5 and Fable 5 finished behind the leaders, showing that broadly correct crisis recognition did not erase differences in execution.
- Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last in the model field.
Opus 4.8’s result is perhaps the most instructive. Its analysis was extensive, but the close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly. Thoroughness, in other words, did not guarantee operational judgment.
There is also an important fairness qualification. Kimi K3 ran without an effort parameter and therefore used the API default, while the other models ran at xhigh. That difference should accompany any interpretation of the close standings.
A Company Under Genuine Operating Pressure
The environment is more than a collection of prompts. The live company has 13 synthetic employees and real money mechanics, burning €105k each month against €2.3k MRR. Its cash countdown is public, it has accumulated 680+ self-learned playbook rules, and every workday is versioned. The experiment is real and watchable, allowing readers to follow management behavior as an ongoing record rather than accept a polished demonstration.

What Older Adults, Families and Care Leaders Should Notice
The quiz makes AI differences visible without requiring technical expertise. A reader can compare decisions and ask practical questions: Which response checked the available evidence? Which protected trust? Which model recognized the correct next step but stopped before completing it? The answers reveal why selecting AI by writing style alone can be misleading.
For senior-care organizations considering AI, the larger lesson is to test behavior in realistic situations before granting responsibility. Firmulate’s enterprise pilot applies the same wargame approach to a read-only export of a business, with nothing writing back to real systems. That separation permits observation without allowing the experiment to alter operational records.
No league table can decide whether an AI belongs in a particular care setting. What this experiment demonstrates is narrower and more useful: models facing identical facts can show measurably different habits. Some investigate more deeply, some maintain cleaner discipline, and some understand the assignment yet fail to close the loop. When people depend on reliable follow-through, that distinction is not cosmetic. It is the heart of the decision.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html