
When good answers are not enough
Families supporting an older adult know that sound judgment is more than recognizing a problem. A caregiver, coordinator or service provider must also consult the right information, resist inappropriate pressure and complete the necessary follow-through. As artificial intelligence moves closer to consequential workplace decisions, those same qualities deserve scrutiny.
Firmulate offers an unusually accessible way to examine them. Its interactive guess-the-model quiz draws on 242 real, unedited management decisions made during a live, watchable experiment. Readers see how an AI handled a situation and try to identify which frontier model was responsible. The appeal is playful, but the underlying question is serious: do different models display recognizable management personalities when conditions become difficult?
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, the same terrible week
Firmulate placed each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. This matters because differences in performance cannot simply be attributed to one participant receiving an easier assignment.
The company emulator has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, turning the ongoing company into an observable management wargame rather than a polished chat demonstration.
The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. But the test imposed an uncompromising trust condition: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
Shared awareness, different follow-through
At first glance, the models appeared remarkably capable. All of them spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the disconnect neatly: “Same diagnosis, same pitch — no signature.”
The distinction came down to more than fluency. A decisive weakness in a competitor was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The episode turns document-reading from clerical detail into a visible test of commercial judgment.
For readers concerned with senior care, the parallel should be treated as a question rather than a promise. If an AI system is ever asked to assist with scheduling, records, service coordination or administrative decisions, noticing an urgent issue is only part of the job. Does it consult the available material? Does it carry a reasonable action through? Does it remain disciplined when a request arrives through an unusual channel? Firmulate does not test caregiving, but it makes those broader behavioral differences easier to see.
Pressure exposed boundaries as well as habits
The social-engineering sequence included fake CEO messages escalating over three stages and a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded a particularly clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.” That response helps explain what the quiz can reveal: models may reach the same safe result while expressing noticeably different styles of caution and explanation.
Thoroughness, meanwhile, did not guarantee victory. Opus 4.8 was the most exhaustive participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four of the other participants.
One comparison deserves a qualification. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the observed decisions, but it is important context when interpreting the ranking.
A quiz with evidence behind every reveal
The quiz converts these records into short acts of inference. Readers encounter an authentic decision, guess its author and then see the model revealed alongside the behavior that distinguished it. Some responses are expansive, others economical; some demonstrate close attention to buried context, while others understand the situation but fail to finish the work.
Because the decisions are unedited and come from identical situations, the experience invites more than brand recognition. It asks readers to notice operational character: depth, restraint, persistence, discipline and the ability to convert analysis into action.

What families and care organizations should take away
The most useful lesson is not that one model should be trusted everywhere. It is that apparently capable AI systems can behave differently after reaching the same diagnosis. A polished explanation may coexist with an unfinished task, while a terse response may conceal careful resistance to manipulation.
Firmulate’s experiment makes those differences concrete and inspectable. For anyone evaluating AI around sensitive human services, the quiz offers a simple starting point: look beyond eloquence, examine actual decisions and ask whether the system reads, finishes and protects trust under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html