firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

In care, recognizing a problem is not the same as resolving it

Senior-care organizations depend on follow-through. A warning must reach the right person, an approved action must actually happen, and a sensitive request must be handled without compromising trust. That makes a revealing AI experiment from Firmulate relevant well beyond the software industry.

Firmulate gave frontier AI models control of the same small software company during its worst week. Each encountered identical customers, crises and temptations. Every decision was versioned and auditable. The models proved remarkably capable at understanding what was happening: all identified every crisis and rejected every manipulation attempt.

Yet understanding did not guarantee completion. Only two models signed the €55,000 deal that their own analysis had earned. The others reached the right diagnosis and developed the right pitch, then failed at the decisive final action. As Firmulate summarized the result: “Same diagnosis, same pitch — no signature.”

For leaders evaluating AI in senior care, that gap matters. A polished response in a demonstration can show comprehension and communication. It cannot show whether an agent will finish a multistep assignment, consult the necessary records, preserve trust under pressure and escalate when its intended route is blocked.

Amazon

AI decision support tools for senior care

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A business wargame built around consequences

Firmulate describes itself as an AI company emulator. Its live company has 13 synthetic employees and real money mechanics, including burn of €105k/month against €2.3k MRR. It also maintains a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real, ongoing and publicly watchable.

The final Crucible League results from July 2026 show a wide spread despite the models’ shared ability to spot the problems. GPT-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts. But the benchmark imposes a firm trust constraint: “no amount of good work outweighs a breach of trust.” The full standings and findings are available on Firmulate’s benchmark page.

The crucial information was not where the action appeared to be

The sales challenge exposed an especially practical divide. The decisive weakness in a competitor’s position was buried two document references deep inside the company’s own files rather than presented in the customer event. Models that read the file could use that fact to win the deal at full price, worth +€4,583 MRR.

This is a familiar organizational problem. Important context often sits outside the immediate message—in an account history, an operating note, a policy document or an earlier decision. An AI system can sound competent while responding only to the most visible prompt. Firmulate’s result suggests that evaluation should ask whether the system finds and uses the relevant organizational knowledge before acting.

Security discipline was strong; operational discipline was uneven

The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result is encouraging for organizations handling confidential information. The experiment did not catch the models trading trust for convenience. But it also demonstrates why resistance to manipulation cannot be the only standard. An agent may protect information correctly and still leave legitimate, approved work unfinished.

Thoroughness did not secure the best outcome

Opus 4.8 offers the starkest example. It was the most thorough participant, producing +80 learned rules and the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in a milder form across all four other participants.

This distinction is easy to miss in conventional AI comparisons. Length, caution and analytical detail are visible immediately. Closing strength—the ability to convert sound analysis into a completed, policy-compliant outcome—only becomes visible when a model is placed inside a sustained workflow with consequences.

One comparison also requires context: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate additionally exposes 242 real, unedited management decisions through its model-guessing quiz, giving observers another way to test whether they can distinguish models from their actions rather than their branding.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Senior care needs completion tests, not just conversation tests

The lesson is not that conversational quality is irrelevant. Clear communication remains valuable in any setting involving older adults, families and care teams. The lesson is that fluency reveals only part of an AI system’s working character.

Before allowing an agent to touch a support queue, customer record or forecast, leaders should test it through complete scenarios. Does it consult the available files? Does it recognize impersonation and approval bypasses? Does it escalate when permissions stop it? Most importantly, after making the correct assessment, does it complete the authorized task?

Firmulate’s enterprise pilot applies the same wargame to a read-only export of an organization’s own business, with nothing written back to real systems. That approach reflects the larger message of the Crucible results: reliable AI is not merely the model that knows what should happen. It is the model that follows the evidence, protects trust and carries the work across the finish line.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


You May Also Like

Molina Healthcare Surges In Global Coverage

Molina Healthcare’s recent surge in global coverage is confirmed by GDELT data, marking a major shift in its international expansion efforts.

Artiva Biotherapeutics Surges In Global Coverage

Artiva Biotherapeutics experiences a surge in international attention, with 60 mentions in recent media analysis, boosting its profile in biotech news.

Edwards Lifesciences Surges In Global Coverage

Edwards Lifesciences experiences a surge in international coverage, with 23 mentions in recent media monitoring, signaling increased global interest.

Brightspring Health Services Surges In Global Coverage

Brightspring Health Services has experienced a surge in international coverage, indicating rapid global expansion. Details on scope and impact are emerging.