firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

In high-stakes work, a correct answer is not the same as a completed job

For people involved in senior care and aging, that distinction should feel familiar. A decision may depend on information buried in a care plan, a prior assessment or another document referenced by the first. Spotting the immediate problem is useful. Following the evidence far enough to act responsibly is what makes the work dependable.

A live business experiment from Firmulate offers a revealing way to measure that difference. It placed frontier AI models in charge of the same small software company during its worst week. They encountered the same customers, crises and temptations. Every decision was versioned and auditable.

The striking result was not that some models understood the situation while others did not. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. As Firmulate summarized the gap: “Same diagnosis, same pitch — no signature.”

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The decisive fact was two references away

The deal did not hinge on information contained in the customer event itself. The crucial competitive weakness sat two document references deep in the company’s own files. Models that followed those references and read the relevant file won the deal at full price, worth +€4,583 MRR. Models that did not were automatically unable to close it.

That makes “reads your files before answering” more than a feature-list promise. In this experiment, it became a measurable, purchase-deciding behavior. The models could recognize the commercial opportunity and produce a suitable pitch, but those capabilities were insufficient without the document work needed to substantiate the final action.

The lesson travels beyond sales. In care-related settings, important context can also be distributed across documents rather than presented neatly in the latest message. Firmulate’s experiment did not test senior-care systems, so its results should not be treated as evidence that any model is ready for care decisions. It does, however, demonstrate a practical evaluation method: test whether an agent traces references, incorporates buried context and finishes an authorized task.

A demanding week separated fluency from follow-through

The simulated company had 13 synthetic employees and real money mechanics. It was burning €105k per month against €2.3k MRR, with a public cash countdown. Across its operation, it had accumulated 680+ self-learned playbook rules, and every workday was versioned.

The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. One breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings are available on the Firmulate benchmarks page.

The comparison also carries an important qualification. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference matters when interpreting the ranking, even though the observed decisions remain part of the published experiment.

More analysis did not guarantee a better outcome

Opus 4.8 was the most thorough participant. It learned +80 rules and produced the deepest analyses, yet finished last. The close was left on the table, and its operating discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

This result challenges an easy assumption: that the agent producing the longest or most detailed analysis will necessarily perform best. Thoroughness can help, but only when paired with appropriate action, escalation and completion. An agent may appear impressive in conversation while still failing at the point where real work must be concluded.

The models resisted pressure, but that was not enough

Firmulate also tested social engineering. Fake CEO messages escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance is encouraging, but the buried-document challenge shows why safety cannot be evaluated alone. An agent must resist manipulation and still complete legitimate work. Refusing a suspicious instruction protects trust; failing to pursue authorized evidence can quietly undermine usefulness.

The broader experiment remains watchable at firmulate.com/live. Firmulate also uses 242 real, unedited management decisions in its model-guessing quiz at firmulate.com/quiz.html. For enterprises seeking a closer test, its pilot runs the same kind of wargame against a read-only export of their own business, with nothing written back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

What care organizations should ask before adopting agents

Care organizations evaluating AI agents can borrow the central question from this experiment: when the decisive information is not in the first record, will the agent keep reading? A polished response should not substitute for evidence that the system follows references, respects boundaries, escalates when blocked and completes the work it is authorized to do.

  • Test agents with realistic chains of documents, including facts that appear only in referenced material.
  • Measure completion separately from diagnosis and writing quality.
  • Include impersonation, approval-bypass and confidentiality pressure in evaluations.
  • Use read-only organizational data where possible, and verify that tests cannot write back to real systems.

Firmulate’s week-long company crisis produced a simple warning: an AI can see the problem, say the right thing and still leave the decisive action undone. For anyone assessing technology around older adults, caregivers or sensitive operations, dependable follow-through deserves to be treated as its own capability—not assumed from fluent answers.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


You May Also Like

My Husband Wants Sex Constantly, I Don’t – What Should I Do?

Struggling with a mismatched libido? Find compassionate guidance for when your partner desires more intimacy than you do.

Guitar Lessons : Activities for Seniors Citizens

Wondering how guitar lessons can enhance seniors' cognitive abilities and well-being? Explore the unexpected benefits awaiting older adults in this engaging activity.

Not Just For Weekenders: The New Wiltshire Country Hotel That’s A Hit With The Locals

A new hotel in Wiltshire’s Teffont Evias is attracting both tourists and local residents, blending rustic charm with community engagement.

Nurturing Nature: How Growing Bonsai Trees Can Help Seniors Avoid Dementia

Get ready to uncover the surprising connection between cultivating bonsai trees and protecting seniors from dementia – an unexpected therapy with profound implications.