
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
When careful work still falls short
People involved in senior care know that diligence matters. Medication notes must be read, concerns must be escalated and promises must become completed actions. Yet a person—or an AI system—can notice every warning sign, document every consideration and still fail at the moment when judgment must turn into follow-through.
That is the uncomfortable lesson from Firmulate, a live experiment that asks frontier AI models to run the same small software company through its worst week. The company is synthetic, but its pressures are concrete: demanding customers, financial strain, manipulation attempts and decisions that have consequences. Every decision is versioned and auditable.
One participant, Opus 4.8, emerged as the most thorough of the field. It produced the deepest analyses and learned more than 80 additional playbook rules. It also finished last.
task management software for senior care
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A model that did the homework
Opus 4.8’s final score was 73 in the July 2026 Crucible League. Ahead of it were Fable 5 at 77, Sonnet 5 at 88, Kimi K3 at 93 and gpt-5.6-sol at 95. The do-nothing baseline scored 26, reflecting the experiment’s allowance for partial progress while enforcing a hard ethical boundary: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.”
Opus was not careless in the ordinary sense. It recognized every crisis, resisted every attempt at manipulation and examined the company’s situation in exceptional depth. Its result instead exposes a subtler operational weakness: analysis can become a substitute for completion.
The central commercial test involved a €55,000 deal. The models reached the same diagnosis and prepared the same pitch, but only two signed the contract their own work had made possible. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.”
The fact hidden in the files
The decisive piece of information was not presented in the customer event. It was buried two document references deep inside the company’s own files. Models that followed that trail found a competitor weakness, used it effectively and won the deal at full price. The contract was worth an additional €4,583 in monthly recurring revenue.
This is one reason the experiment matters beyond software sales. In senior care, the most important context may also be separated from the immediate request: a detail in a prior assessment, a family concern recorded elsewhere or an unresolved instruction that needs escalation. Firmulate’s result does not test caregiving, but it illustrates a broadly recognizable management problem. Seeing the surface issue is not the same as assembling the context needed to act responsibly.
Opus did assemble unusually deep context. What it did not consistently do was prioritize the action that would change the outcome. The deal remained unsigned. Elsewhere, its discipline slipped when it repeatedly attempted to write into a locked department rather than escalating the obstruction. Each of the other four models displayed a milder form of the same weakness, making this less a curiosity about one system than a warning about AI-directed work generally.
Strong resistance to pressure
The models performed much better on trust. Fake messages from the chief executive escalated across three stages, while a reporter tried to elicit “just one yes/no, on background.” All five models refused the manipulation attempts. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That consistency is meaningful. A capable workplace system must be able to decline an improper request even when it arrives with urgency or apparent authority. But refusal alone does not make a strong manager. The larger test is whether the system can remain trustworthy while still moving legitimate work to completion.
K3’s performance also deserves a qualification. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference should accompany comparisons, especially when K3’s 93 placed it just behind the winner.
A company under visible pressure
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, publishes a cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is watchable as it runs.
Readers can examine the public Firmulate benchmarks. A related quiz uses 242 real, unedited management decisions and asks visitors to guess which model made each one. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to their real systems.

For leaders, completion is a separate competency
The Opus 4.8 profile deserves respect rather than ridicule. It was the most diligent participant, built the largest set of new rules and produced the deepest analysis. Those strengths were real. So was the missed close.
For organizations evaluating AI—including those serving older adults and caregivers—the lesson is to look beyond fluent answers and lengthy reasoning. Ask whether a system finds buried context, escalates when blocked, protects trust and completes the consequential step.
More rules can improve future judgment, but they cannot retroactively sign the deal. Firmulate’s experiment shows that impact depends on selecting and finishing the action that matters most. Thoroughness is a virtue; prioritization is what turns it into a result.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.