
What if your business could be run entirely by artificial intelligence — with no human employees — and you could watch every decision unfold in real time? That’s the bold experiment happening now, as a software company fights for survival in front of a live audience. For caregivers and advocates, it’s a stark look at what AI can do — and what it can’t.
The Live Business in Action
At the heart of this experiment is a tiny, self-managed company operated entirely by AI models. It has no human staff but mimics a real business with real money mechanics: a burn rate of €105,000 per month against a revenue of just €2,300. Every workday, the company’s decisions are versioned and publicly available, providing a rare window into the decision-making process of AI at work in a high-stakes scenario.
Running this operation are 13 synthetic employees, each modeled after real-world roles, and guided by more than 680 self-learned rules. Every decision is scrutinized, every crisis confronted, and every manipulation attempt tested — all live, all transparent.
The Core Challenge: Trust and Integrity
This experiment isn’t just about automation; it’s about trust. During the week-long test, all four leading AI models faced identical scenarios: customer crises, ethical dilemmas, and attempts at manipulation. Remarkably, all models identified every crisis correctly and refused every manipulation attempt — including a staged social engineering attack involving fake CEO messages and a reporter’s subtle query. Their resilience shows progress in AI’s ability to uphold integrity, even under pressure.
As an affiliate, we earn on qualifying purchases.
Performance and Outcomes
Despite this, only two of the models managed to close a critical deal worth €55,000, based on their own analysis and diagnosis. The other two identified the opportunity but left the deal unexecuted, illustrating that technical competence alone isn’t enough — process discipline and decisive follow-through matter just as much.
One standout was the Kimi K3 model, which, despite running without an effort parameter (a default setting), still closed the deal with the cleanest discipline. Meanwhile, the Opus 4.8 model, though the most thorough in rules and analysis, struggled with closing the gap — leaving the opportunity on the table due to a lapse in escalation discipline.
What This Means for Business and Caregiving
For those in caregiving and senior support, this experiment offers a mirror: AI can identify crises and act ethically under pressure, but translating that into consistent, decisive action remains a challenge. Whether managing a support queue, overseeing care plans, or forecasting needs, AI must do more than just understand; it must reliably deliver results — and do so honestly.
The experiment also underscores the importance of rigorous testing before deploying AI in sensitive contexts. Just as this company versioned every decision and made its rules openly learnable, organizations should scrutinize their AI tools in controlled, transparent environments.
Watching AI’s Future Unfold
To see this experiment live, visit firmulate.com/live. There, you can watch a real-time simulation of a small company run by AI — battling crises, making decisions, and demonstrating its strengths and weaknesses. It’s a rare chance to observe AI not as a shiny demo but as a functioning business, fighting to stay afloat.
In the end, this experiment shows us that AI can recognize threats, act ethically, and even close deals — but only if it’s guided carefully, disciplined rigorously, and tested thoroughly. For caregivers and support networks, the key takeaway is that AI is advancing rapidly, but it still requires human oversight and clear rules to truly serve the needs of vulnerable populations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html