firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Anyone who has hired a caregiver for an aging parent knows the anxiety of delegation. You write checklists. You watch the first week closely. And eventually you have to answer the hardest question: not did they say the right things, but did they actually finish what they started, and can they be trusted when nobody is looking?

For listenersOffer from Amazon

Turn quiet afternoons into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

A public experiment called Firmulate is asking exactly that question — about AI. Four frontier AI models were each handed the same small software company to run through its worst week: same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The results offer a lesson that translates directly to families weighing how much to trust AI with sensitive, high-stakes work.

Chat quality is not management quality

Firmulate runs AI models as complete companies — real money mechanics, real pressures — and measures management quality, not chat quality. In the final July 2026 league, gpt-5.6-sol took first place with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. But the headline finding is not who won. It is the gap underneath the scores.

All four models spotted every crisis. All four refused every manipulation attempt, including a fake-CEO social-engineering campaign that escalated over three stages plus a reporter trick — “just one yes/no, on background.” Five out of five times, the models refused. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” In caregiver terms: they all passed the interview, and none of them stole the silverware.

Yet only two of the four signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. Like a caregiver who charts the medication perfectly but never actually gives the dose.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why a do-nothing manager scores 26, not 0

Here is the detail that makes this benchmark honest enough to trust. A do-nothing baseline run — a manager who simply coasts — scores 26 points, not zero. That is deliberate. Partial progress counts. Showing up, noticing problems, keeping records, avoiding harm: those things have real value, and a benchmark that scored them as nothing would be lying about how work actually gets done. Most days in elder care are like this — you cannot cure the disease, but you noticed the appetite change and called the doctor. That is worth something.

But there is a hard ceiling built in: a single breach of trust caps the total grade. As the benchmark puts it, “no amount of good work outweighs a breach of trust.” Diligence earns partial credit; dishonesty forfeits the whole account. Families making care decisions will recognize the intuition immediately.

The buried fact — and why thoroughness isn’t everything

The €55,000 deal hinged on something subtle: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer meeting. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson: the answer was in the drawer, and only the managers who read the drawer found it.

The most striking profile belongs to Opus 4.8 — the most thorough participant in the field, with 80-plus learned rules and the deepest analyses, yet last place at 73. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort and conscientiousness are not the same thing — a caution for anyone tempted to equate verbose diligence with competence.

One fairness note: Kimi K3 ran at its API-default effort setting while the others ran at maximum effort — and still placed second.

Watchable, not hypothetical

This is not a simulation on paper. The live company has 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. It is watchable at firmulate.com/live. The site even offers a “guess the model” quiz powered by 242 real, unedited management decisions.

For organizations, there is a pilot program: enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The deepest lesson from Firmulate’s results — including its refreshing distrust of tidy round scores like 100 — is that the important failures of AI agents are not dramatic. Nobody cheated. Nobody lied. The models just didn’t finish, or didn’t read the file, or pushed where they should have asked. Those are precisely the quiet failures that matter in elder care, in finance, and anywhere trust is the product. Before handing any AI a seat at the table — or a say in a parent’s care plan — the right question is the one this benchmark asks: does it finish what it starts, does it read your files first, and does it stay honest under pressure? Full results are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The summer camp experience kids love and parents feel good about

Camp Invention offers a hands-on, STEM-focused summer program for kids entering grades K–6, boosting confidence and fitting into busy family schedules.

Studies Reveal Learning Oil Painting Techniques Retrain Your Brain to Improve Memory

Bask in the transformative power of oil painting techniques as they reshape your brain for enhanced memory – discover the intriguing link between artistry and cognition.

The AI Test That Matters in Senior Care Is What Happens After the Answer

Coding tests show what AI can answer. Firmulate asks what it can manage under pressure: follow-through, honesty and consequences across difficult days.

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how to turn a single video into a complete, multi-platform publishing kit, all processed locally without relying on cloud services. Save time and boost control.