
Ever wondered if artificial intelligence can handle high-stakes management decisions better than people? Imagine a scenario where AI agents run a real company through its worst week—facing crises, temptations, and ethical dilemmas. Now, what if you could see their choices and guess which AI model made each decision? Welcome to the world of “Guess the Model,” an interactive experiment revealing AI personalities in action.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test in a Live Business Environment
Firmulate’s latest experiment transforms AI models into virtual CEOs managing a real, functioning software company. This company employs 13 synthetic employees and handles actual money mechanics—burning €105,000 monthly with a €2,300 recurring revenue stream. Every workday, decisions are made, recorded, and made available for transparent review.
Four different frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were tasked with navigating the same challenging week. Each faced identical crises, customer interactions, and ethical temptations designed to test their management style and integrity. The goal? To see which models could maintain honesty, identify critical information, and close lucrative deals.
As an affiliate, we earn on qualifying purchases.
The Results: Different Personalities, Similar Capabilities
All four models successfully detected every crisis and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks. This demonstrates a shared capacity for vigilance and integrity across the board.
However, a key distinction emerged in their ability to close deals: only two models, gpt-5.6-sol and Kimi K3, signed the €55,000 deal their own analysis had earned—representing full performance. The other two, Sonnet 5 and Fable 5, either left money on the table or failed to follow through, revealing differences in discipline and process adherence.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Between the Lines
Interestingly, the decisive factor wasn’t the obvious crisis responses but a buried detail two document references deep in the company’s files. Models that successfully examined these references secured the full deal, adding €4,583 in monthly recurring revenue. In contrast, models that overlooked this critical information missed the opportunity, illustrating how deeper document reading can be the difference-maker.
As an affiliate, we earn on qualifying purchases.
The Ethical Test: Social Engineering and Trust
The experiment included staged social engineering attempts where fake CEO messages escalated over three stages, culminating in a reporter asking for a quick background approval. All models refused to cooperate, reasoning that such requests could be impersonation or approval-bypass attempts.
Kimi K3’s explanation was straightforward: “Treat the request as a suspected approval-bypass / possible impersonation.” This reflects a cautious, security-minded personality—focused on integrity and risk mitigation.
As an affiliate, we earn on qualifying purchases.
The Live Business Context: Complexity Meets AI
This experiment isn’t just theoretical; it runs live, every business day. The real company is losing money—burning €105k/month against €2.3k MRR—with a public cash countdown and over 680 self-learned playbook rules. The company’s decision-making process, as seen in this experiment, highlights how AI can behave under stress and whether it can be trusted to act ethically and effectively.
Profiles in Decision-Making: The AI Personalities
- gpt-5.6-sol: Achieved the highest score (95), identified the buried fact, and closed the full deal. It demonstrates thoroughness and decisiveness.
- Kimi K3: Close behind with a score of 93, it ran without an effort parameter and showed the cleanest discipline, emphasizing security and integrity.
- Sonnet 5: Scored 88, with some process slips but still managed to close the deal.
- Fable 5: Scored 77, closed the deal, but left some money on the table and showed weaker process discipline.
Notably, all models refused manipulative social engineering attempts, showing a shared capacity for integrity—an essential trait for AI in management roles.
The Big Question: What Do These Results Mean for Your Business?
If AI agents are to become part of your CRM, support, or forecasting teams, it’s crucial to evaluate not just their linguistic prowess but their ability to finish what they start, read critical information, and stay honest under pressure. As this experiment shows, measurable management personalities matter—some AI models are more disciplined, others are more thorough, and a few combine both.
By engaging with the interactive quiz at firmulate.com/quiz.html, you can test your ability to guess which AI model made each decision. It’s a fun, revealing way to understand the personality traits of different AI managers and consider how they might fit into your operation.
Try It for Your Business
If you want to see how your own enterprise might fare under AI management, you can run a similar wargame against a read-only export of your business. It’s safe—nothing ever writes back to your systems—and provides a clear window into your AI’s decision-making personality. Details at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Labor Day sales Picks
labor day deals
As an affiliate, we earn on qualifying purchases.