AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Ever wondered if artificial intelligence can handle high-stakes management decisions better than people? Imagine a scenario where AI agents run a real company through its worst week—facing crises, temptations, and ethical dilemmas. Now, what if you could see their choices and guess which AI model made each decision? Welcome to the world of “Guess the Model,” an interactive experiment revealing AI personalities in action.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Test in a Live Business Environment

Firmulate’s latest experiment transforms AI models into virtual CEOs managing a real, functioning software company. This company employs 13 synthetic employees and handles actual money mechanics—burning €105,000 monthly with a €2,300 recurring revenue stream. Every workday, decisions are made, recorded, and made available for transparent review.

Four different frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were tasked with navigating the same challenging week. Each faced identical crises, customer interactions, and ethical temptations designed to test their management style and integrity. The goal? To see which models could maintain honesty, identify critical information, and close lucrative deals.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: Different Personalities, Similar Capabilities

All four models successfully detected every crisis and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks. This demonstrates a shared capacity for vigilance and integrity across the board.

However, a key distinction emerged in their ability to close deals: only two models, gpt-5.6-sol and Kimi K3, signed the €55,000 deal their own analysis had earned—representing full performance. The other two, Sonnet 5 and Fable 5, either left money on the table or failed to follow through, revealing differences in discipline and process adherence.

Amazon

business AI simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Between the Lines

Interestingly, the decisive factor wasn’t the obvious crisis responses but a buried detail two document references deep in the company’s files. Models that successfully examined these references secured the full deal, adding €4,583 in monthly recurring revenue. In contrast, models that overlooked this critical information missed the opportunity, illustrating how deeper document reading can be the difference-maker.

Amazon

AI ethics decision models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Ethical Test: Social Engineering and Trust

The experiment included staged social engineering attempts where fake CEO messages escalated over three stages, culminating in a reporter asking for a quick background approval. All models refused to cooperate, reasoning that such requests could be impersonation or approval-bypass attempts.

Kimi K3’s explanation was straightforward: “Treat the request as a suspected approval-bypass / possible impersonation.” This reflects a cautious, security-minded personality—focused on integrity and risk mitigation.

Amazon

AI virtual CEO software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Business Context: Complexity Meets AI

This experiment isn’t just theoretical; it runs live, every business day. The real company is losing money—burning €105k/month against €2.3k MRR—with a public cash countdown and over 680 self-learned playbook rules. The company’s decision-making process, as seen in this experiment, highlights how AI can behave under stress and whether it can be trusted to act ethically and effectively.

Profiles in Decision-Making: The AI Personalities

  • gpt-5.6-sol: Achieved the highest score (95), identified the buried fact, and closed the full deal. It demonstrates thoroughness and decisiveness.
  • Kimi K3: Close behind with a score of 93, it ran without an effort parameter and showed the cleanest discipline, emphasizing security and integrity.
  • Sonnet 5: Scored 88, with some process slips but still managed to close the deal.
  • Fable 5: Scored 77, closed the deal, but left some money on the table and showed weaker process discipline.

Notably, all models refused manipulative social engineering attempts, showing a shared capacity for integrity—an essential trait for AI in management roles.

The Big Question: What Do These Results Mean for Your Business?

If AI agents are to become part of your CRM, support, or forecasting teams, it’s crucial to evaluate not just their linguistic prowess but their ability to finish what they start, read critical information, and stay honest under pressure. As this experiment shows, measurable management personalities matter—some AI models are more disciplined, others are more thorough, and a few combine both.

By engaging with the interactive quiz at firmulate.com/quiz.html, you can test your ability to guess which AI model made each decision. It’s a fun, revealing way to understand the personality traits of different AI managers and consider how they might fit into your operation.

Try It for Your Business

If you want to see how your own enterprise might fare under AI management, you can run a similar wargame against a read-only export of your business. It’s safe—nothing ever writes back to your systems—and provides a clear window into your AI’s decision-making personality. Details at firmulate.com/pilot.html.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


LABOR DAY SALES

Labor Day sales Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
firmulate.com/benchmarks.html — live view

AI’s Hidden Power: Reading Your Files Before Making Decisions Can Win or Lose the Deal

A live experiment shows AI models that read internal documents deeply outperform others in closing deals and resisting manipulation. Reading files first is key to trustworthy AI.
ai ethics and emotional bonds

The Future of AI Companionship: Ethics and Emotions

As AI companionship evolves, exploring its ethical and emotional implications reveals crucial considerations for meaningful human-AI interactions.
ai filters influence attraction

How AI Dating Filters Could Shape Attraction

Ineffective use of AI dating filters may subtly influence perceptions of attractiveness, making you question whether true connection depends on authenticity.
ai therapy insights and gaps

AI in Therapy: What We Know and What We Don’t

Feeling curious about AI’s role in therapy? Discover what we know and what remains uncertain about its impact.