
Imagine a bustling music studio where every decision, every note, is watched and audited in real-time. Now, picture a software company that operates without employees, losing money daily, and yet, viewers can observe its every move. This isn’t science fiction — it’s the stark reality of the live experiment by Firmulate, where AI models run a virtual company in real-world crises, revealing not just how AI performs, but how honest and disciplined it can be under pressure.
The Living Laboratory: Watching a Company Fight for Survival
At the heart of this experiment is a small, synthetic software company managed solely by AI models. Instead of human employees, there are 13 simulated workers executing tasks based on over 680 learned rules, with every decision versioned and published daily. The company operates under real financial constraints: burning €105,000 each month against a modest €2,300 monthly recurring revenue. It’s a high-stakes environment designed to test AI’s management skills amidst crises, temptations, and strategic decisions. Viewers can explore this ongoing story at firmulate.com/live.html.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How AI Models Tackle Crises and Manipulation
The experiment pits four frontier AI models against the same challenging week. Each is tasked with navigating customer crises, avoiding manipulative requests, and closing deals. Remarkably, all four models identified every crisis, refused every manipulation attempt, and maintained integrity. The only hitch? Only two of them managed to close a €55,000 deal, which their own analyses justified. The other two recognized opportunities but did not act, leaving potential revenue unrealized. This highlights a critical insight: AI can recognize opportunities and threats just as well as humans, but execution depends on discipline and process adherence.
The Hidden Weakness: Knowledge in Files
Deep within the company’s files lay a crucial piece of information that could have sealed the deal at full price (+€4,583 MRR). Yet, the AI models that read only superficial data missed this key detail, while those that examined deeper references succeeded. This underscores an essential lesson: the quality of a model’s decision-making heavily depends on its ability to access and interpret relevant internal data, not just surface information or external cues.
Dealing with Social Engineering and Trust
Another test involved simulated social engineering attacks—fake CEO messages escalating in seriousness and a reporter trick. All five models refused to act on these requests, with Kimi K3 explicitly noting that such requests could be impersonation attempts. This demonstrates that AI systems can be designed with built-in skepticism, refusing manipulative or suspicious prompts, a crucial feature for any AI operating in sensitive business environments.
The Reality of a Money-Losing Company
Despite the sophisticated decision-making, the company is a financial sinkhole, burning €105,000 monthly with only €2,300 in recurring revenue. This stark reality is visible in the public dashboard, emphasizing the ‘build-in-public’ ethos of the experiment. It’s a raw, unfiltered view into the harsh conditions under which AI management must operate, providing invaluable insights for organizations considering AI-driven workflows in customer support, sales, or operations.
Lessons for Business and AI Developers
The experiment’s most profound takeaway is that AI can recognize complex crises and resist manipulation, but execution and discipline remain challenging. For example, Opus 4.8, with the most thorough analysis, finished last because it left a potential deal on the table, demonstrating discipline lapses. Similarly, models defaulted to safe, conservative decisions in the face of social engineering, showcasing their built-in safeguards.
Take Action: Run Your Own AI Wargame
Businesses can now run similar simulations against their own operations, testing how AI would handle crises, temptations, and strategic decisions without risking real-world consequences. These tests are accessible through firmulate.com/pilot.html and offer a safe environment to evaluate AI readiness before deployment. The experiment is ongoing, with new runs and insights published regularly, making it a living resource for anyone interested in AI’s practical capabilities in management.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html