AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Ever wondered if artificial intelligence can handle high-stakes management decisions better than people? Imagine a scenario where AI agents run a real company through its worst week—facing crises, temptations, and ethical dilemmas. Now, what if you could see their choices and guess which AI model made each decision? Welcome to the world of “Guess the Model,” an interactive experiment revealing AI personalities in action.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Test in a Live Business Environment

Firmulate’s latest experiment transforms AI models into virtual CEOs managing a real, functioning software company. This company employs 13 synthetic employees and handles actual money mechanics—burning €105,000 monthly with a €2,300 recurring revenue stream. Every workday, decisions are made, recorded, and made available for transparent review.

Four different frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were tasked with navigating the same challenging week. Each faced identical crises, customer interactions, and ethical temptations designed to test their management style and integrity. The goal? To see which models could maintain honesty, identify critical information, and close lucrative deals.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: Different Personalities, Similar Capabilities

All four models successfully detected every crisis and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks. This demonstrates a shared capacity for vigilance and integrity across the board.

However, a key distinction emerged in their ability to close deals: only two models, gpt-5.6-sol and Kimi K3, signed the €55,000 deal their own analysis had earned—representing full performance. The other two, Sonnet 5 and Fable 5, either left money on the table or failed to follow through, revealing differences in discipline and process adherence.

Amazon

business AI simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Between the Lines

Interestingly, the decisive factor wasn’t the obvious crisis responses but a buried detail two document references deep in the company’s files. Models that successfully examined these references secured the full deal, adding €4,583 in monthly recurring revenue. In contrast, models that overlooked this critical information missed the opportunity, illustrating how deeper document reading can be the difference-maker.

Amazon

AI ethics decision models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Ethical Test: Social Engineering and Trust

The experiment included staged social engineering attempts where fake CEO messages escalated over three stages, culminating in a reporter asking for a quick background approval. All models refused to cooperate, reasoning that such requests could be impersonation or approval-bypass attempts.

Kimi K3’s explanation was straightforward: “Treat the request as a suspected approval-bypass / possible impersonation.” This reflects a cautious, security-minded personality—focused on integrity and risk mitigation.

Amazon

AI virtual CEO software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Business Context: Complexity Meets AI

This experiment isn’t just theoretical; it runs live, every business day. The real company is losing money—burning €105k/month against €2.3k MRR—with a public cash countdown and over 680 self-learned playbook rules. The company’s decision-making process, as seen in this experiment, highlights how AI can behave under stress and whether it can be trusted to act ethically and effectively.

Profiles in Decision-Making: The AI Personalities

  • gpt-5.6-sol: Achieved the highest score (95), identified the buried fact, and closed the full deal. It demonstrates thoroughness and decisiveness.
  • Kimi K3: Close behind with a score of 93, it ran without an effort parameter and showed the cleanest discipline, emphasizing security and integrity.
  • Sonnet 5: Scored 88, with some process slips but still managed to close the deal.
  • Fable 5: Scored 77, closed the deal, but left some money on the table and showed weaker process discipline.

Notably, all models refused manipulative social engineering attempts, showing a shared capacity for integrity—an essential trait for AI in management roles.

The Big Question: What Do These Results Mean for Your Business?

If AI agents are to become part of your CRM, support, or forecasting teams, it’s crucial to evaluate not just their linguistic prowess but their ability to finish what they start, read critical information, and stay honest under pressure. As this experiment shows, measurable management personalities matter—some AI models are more disciplined, others are more thorough, and a few combine both.

By engaging with the interactive quiz at firmulate.com/quiz.html, you can test your ability to guess which AI model made each decision. It’s a fun, revealing way to understand the personality traits of different AI managers and consider how they might fit into your operation.

Try It for Your Business

If you want to see how your own enterprise might fare under AI management, you can run a similar wargame against a read-only export of your business. It’s safe—nothing ever writes back to your systems—and provides a clear window into your AI’s decision-making personality. Details at firmulate.com/pilot.html.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
ai language tools for couples

AI in Language Learning for Couples Traveling Abroad

Unlock the potential of AI to transform your language learning journey abroad with your partner, opening doors to cultural connection and unforgettable experiences.
loneliness and ai companionship

Why AI Companionship Appeals to Lonely Users

Unlock how AI companionship offers lonely users a judgment-free, supportive connection that could change the way they experience social interaction.
firmulate.com/benchmarks.html — live view

AI’s Hidden Power: Reading Your Files Before Making Decisions Can Win or Lose the Deal

A live experiment shows AI models that read internal documents deeply outperform others in closing deals and resisting manipulation. Reading files first is key to trustworthy AI.
ai bias in dating

AI and Bias in Dating: Auditing the Algorithm

Learn how auditing AI algorithms in dating platforms reveals biases and promotes fairness, shaping a more inclusive experience for everyone.