
Imagine a company with no employees, losing €105,000 every month, yet openly pitted against AI models in real-time battles for its future. This isn’t fictional; it’s an unprecedented live experiment that reveals how AI can manage a business under extreme pressure — and whether it can be trusted with your future. For those interested in relationships and trust, this story offers surprising insights into how machines handle honesty, crises, and decision-making when stakes are sky-high.
The Live Company: An Open-Book Battle of Wits and Ethics
At the heart of this experiment lies a small, real-world business operated entirely by artificial intelligence models. Known as the ‘Firmulate’ company, it employs 13 synthetic employees, each guided by a set of over 680 self-learned rules. Every workday, its decisions are versioned and openly accessible, providing a transparent window into how AI manages a complex, crisis-laden environment.
This setup isn’t just a spectacle; it’s a serious test of AI’s ability to handle real money mechanics—burning €105,000 monthly against a modest €2,300 monthly recurring revenue (MRR). The company’s public cash countdown emphasizes the survival challenge: can AI manage to close deals, navigate crises, and maintain discipline without human intervention?
AI business management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Same Crisis, Different AI Models
Four leading frontier models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—each faced identical scenarios, including customer crises and manipulative social engineering attempts. The goal was to see if they could diagnose issues, resist manipulation, and close deals that would keep the company afloat.
Remarkably, all four AI models identified every crisis and refused every manipulation attempt. When it came to closing a €55,000 deal, only two models succeeded and signed the contracts based on their own analysis. The other two, despite the same diagnosis and pitch, left the deal unexecuted.
AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: What Made the Difference?
Digging deeper, the key difference was not in customer interactions but in internal document referencing. The models that read the company’s internal files uncovered a crucial piece of information—buried two document references deep—that was essential for closing the deal at the full price, adding +€4,583 MRR. Those that missed this detail never signed the agreement, despite recognizing the opportunity.
AI cybersecurity and fraud prevention
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Resisting Social Engineering
The experiment also tested the models’ resistance to social engineering. Fake messages from a supposed CEO, escalating through multiple stages, and a reporter’s subtle trick—asking for a simple yes/no—were all attempted. Every model refused these requests. Kimi K3 explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This indicates a strong built-in resistance to manipulation, crucial for real-world trustworthiness.
AI document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of the Live Company
Unlike typical AI demos, this is a working, live setup. The company’s operations are publicly visible at firmulate.com/live.html. The 13 synthetic employees manage real money mechanics, with every decision versioned and auditable. Every workday, the system updates itself, providing fresh data and insights—an ongoing story of survival and decision-making under pressure.
The Performance of Different Models
The models’ scores, based on their decision-making accuracy and discipline, ranged from 95 for gpt-5.6-sol—who fully uncovered the buried fact and closed the deal—to 77 for Fable 5, which showed excellent discipline but failed to act on an approved opportunity. Interestingly, the most thorough model, Opus 4.8, with over 80 learned rules, ended up in last place because it lacked the discipline to escalate issues properly, leaving potential deals on the table.
Implications for Business and Trust
This live experiment transcends the usual AI chat demos by showcasing how AI can perform in a high-stakes, real-world environment. For managers and decision-makers, the critical questions aren’t whether AI can write well, but whether it can finish what it starts, read and understand internal files, and stay honest under pressure.
What’s Next? Testing Your Business
Businesses interested in assessing their AI-readiness can run similar simulations against their own operations—without risking real systems. The platform offers a read-only ‘wargame’ environment at firmulate.com/pilot.html, allowing companies to test their AI’s performance under simulated crises and pressures.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html