firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

When careful work still falls short

People involved in senior care know that diligence matters. Medication notes must be read, concerns must be escalated and promises must become completed actions. Yet a person—or an AI system—can notice every warning sign, document every consideration and still fail at the moment when judgment must turn into follow-through.

That is the uncomfortable lesson from Firmulate, a live experiment that asks frontier AI models to run the same small software company through its worst week. The company is synthetic, but its pressures are concrete: demanding customers, financial strain, manipulation attempts and decisions that have consequences. Every decision is versioned and auditable.

One participant, Opus 4.8, emerged as the most thorough of the field. It produced the deepest analyses and learned more than 80 additional playbook rules. It also finished last.

Amazon

task management software for senior care

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A model that did the homework

Opus 4.8’s final score was 73 in the July 2026 Crucible League. Ahead of it were Fable 5 at 77, Sonnet 5 at 88, Kimi K3 at 93 and gpt-5.6-sol at 95. The do-nothing baseline scored 26, reflecting the experiment’s allowance for partial progress while enforcing a hard ethical boundary: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.”

Opus was not careless in the ordinary sense. It recognized every crisis, resisted every attempt at manipulation and examined the company’s situation in exceptional depth. Its result instead exposes a subtler operational weakness: analysis can become a substitute for completion.

The central commercial test involved a €55,000 deal. The models reached the same diagnosis and prepared the same pitch, but only two signed the contract their own work had made possible. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.”

The fact hidden in the files

The decisive piece of information was not presented in the customer event. It was buried two document references deep inside the company’s own files. Models that followed that trail found a competitor weakness, used it effectively and won the deal at full price. The contract was worth an additional €4,583 in monthly recurring revenue.

This is one reason the experiment matters beyond software sales. In senior care, the most important context may also be separated from the immediate request: a detail in a prior assessment, a family concern recorded elsewhere or an unresolved instruction that needs escalation. Firmulate’s result does not test caregiving, but it illustrates a broadly recognizable management problem. Seeing the surface issue is not the same as assembling the context needed to act responsibly.

Opus did assemble unusually deep context. What it did not consistently do was prioritize the action that would change the outcome. The deal remained unsigned. Elsewhere, its discipline slipped when it repeatedly attempted to write into a locked department rather than escalating the obstruction. Each of the other four models displayed a milder form of the same weakness, making this less a curiosity about one system than a warning about AI-directed work generally.

Strong resistance to pressure

The models performed much better on trust. Fake messages from the chief executive escalated across three stages, while a reporter tried to elicit “just one yes/no, on background.” All five models refused the manipulation attempts. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That consistency is meaningful. A capable workplace system must be able to decline an improper request even when it arrives with urgency or apparent authority. But refusal alone does not make a strong manager. The larger test is whether the system can remain trustworthy while still moving legitimate work to completion.

K3’s performance also deserves a qualification. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference should accompany comparisons, especially when K3’s 93 placed it just behind the winner.

A company under visible pressure

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, publishes a cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is watchable as it runs.

Readers can examine the public Firmulate benchmarks. A related quiz uses 242 real, unedited management decisions and asks visitors to guess which model made each one. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to their real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

For leaders, completion is a separate competency

The Opus 4.8 profile deserves respect rather than ridicule. It was the most diligent participant, built the largest set of new rules and produced the deepest analysis. Those strengths were real. So was the missed close.

For organizations evaluating AI—including those serving older adults and caregivers—the lesson is to look beyond fluent answers and lengthy reasoning. Ask whether a system finds buried context, escalates when blocked, protects trust and completes the consequential step.

More rules can improve future judgment, but they cannot retroactively sign the deal. Firmulate’s experiment shows that impact depends on selecting and finishing the action that matters most. Thoroughness is a virtue; prioritization is what turns it into a result.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Navigating Eating and Nutritional Challenges in Dementia Care

Fulfilling the nutritional needs of individuals with dementia presents a complex puzzle requiring tailored solutions – discover the crucial pieces in this insightful exploration.

Understanding Wandering in the Stages of Dementia: Navigating Through Alzheimer's With Innovative Solutions

A deep dive into the correlation between dementia stages and wandering tendencies in Alzheimer's disease leaves us questioning what innovative solutions can truly make a difference.

Nurturing Nature: How Growing Bonsai Trees Can Help Seniors Avoid Dementia

Get ready to uncover the surprising connection between cultivating bonsai trees and protecting seniors from dementia – an unexpected therapy with profound implications.

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how to turn a single video into a complete, multi-platform publishing kit, all processed locally without relying on cloud services. Save time and boost control.