AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Ever wondered if artificial intelligence can handle high-stakes management decisions better than people? Imagine a scenario where AI agents run a real company through its worst week—facing crises, temptations, and ethical dilemmas. Now, what if you could see their choices and guess which AI model made each decision? Welcome to the world of “Guess the Model,” an interactive experiment revealing AI personalities in action.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Test in a Live Business Environment

Firmulate’s latest experiment transforms AI models into virtual CEOs managing a real, functioning software company. This company employs 13 synthetic employees and handles actual money mechanics—burning €105,000 monthly with a €2,300 recurring revenue stream. Every workday, decisions are made, recorded, and made available for transparent review.

Four different frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were tasked with navigating the same challenging week. Each faced identical crises, customer interactions, and ethical temptations designed to test their management style and integrity. The goal? To see which models could maintain honesty, identify critical information, and close lucrative deals.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: Different Personalities, Similar Capabilities

All four models successfully detected every crisis and refused every manipulation attempt, including social engineering tactics like fake CEO messages and reporter tricks. This demonstrates a shared capacity for vigilance and integrity across the board.

However, a key distinction emerged in their ability to close deals: only two models, gpt-5.6-sol and Kimi K3, signed the €55,000 deal their own analysis had earned—representing full performance. The other two, Sonnet 5 and Fable 5, either left money on the table or failed to follow through, revealing differences in discipline and process adherence.

Amazon

business AI simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: Reading Between the Lines

Interestingly, the decisive factor wasn’t the obvious crisis responses but a buried detail two document references deep in the company’s files. Models that successfully examined these references secured the full deal, adding €4,583 in monthly recurring revenue. In contrast, models that overlooked this critical information missed the opportunity, illustrating how deeper document reading can be the difference-maker.

Amazon

AI ethics decision models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Ethical Test: Social Engineering and Trust

The experiment included staged social engineering attempts where fake CEO messages escalated over three stages, culminating in a reporter asking for a quick background approval. All models refused to cooperate, reasoning that such requests could be impersonation or approval-bypass attempts.

Kimi K3’s explanation was straightforward: “Treat the request as a suspected approval-bypass / possible impersonation.” This reflects a cautious, security-minded personality—focused on integrity and risk mitigation.

Amazon

AI virtual CEO software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Business Context: Complexity Meets AI

This experiment isn’t just theoretical; it runs live, every business day. The real company is losing money—burning €105k/month against €2.3k MRR—with a public cash countdown and over 680 self-learned playbook rules. The company’s decision-making process, as seen in this experiment, highlights how AI can behave under stress and whether it can be trusted to act ethically and effectively.

Profiles in Decision-Making: The AI Personalities

  • gpt-5.6-sol: Achieved the highest score (95), identified the buried fact, and closed the full deal. It demonstrates thoroughness and decisiveness.
  • Kimi K3: Close behind with a score of 93, it ran without an effort parameter and showed the cleanest discipline, emphasizing security and integrity.
  • Sonnet 5: Scored 88, with some process slips but still managed to close the deal.
  • Fable 5: Scored 77, closed the deal, but left some money on the table and showed weaker process discipline.

Notably, all models refused manipulative social engineering attempts, showing a shared capacity for integrity—an essential trait for AI in management roles.

The Big Question: What Do These Results Mean for Your Business?

If AI agents are to become part of your CRM, support, or forecasting teams, it’s crucial to evaluate not just their linguistic prowess but their ability to finish what they start, read critical information, and stay honest under pressure. As this experiment shows, measurable management personalities matter—some AI models are more disciplined, others are more thorough, and a few combine both.

By engaging with the interactive quiz at firmulate.com/quiz.html, you can test your ability to guess which AI model made each decision. It’s a fun, revealing way to understand the personality traits of different AI managers and consider how they might fit into your operation.

Try It for Your Business

If you want to see how your own enterprise might fare under AI management, you can run a similar wargame against a read-only export of your business. It’s safe—nothing ever writes back to your systems—and provides a clear window into your AI’s decision-making personality. Details at firmulate.com/pilot.html.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


LABOR DAY SALES

Labor Day sales Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
ai versus human companionship

AI Companions Vs Human Companions: a Comparative Guide

Navigating the choice between AI and human companions reveals surprising insights that could reshape your understanding of connection and support.
How AI Could Replace Animal Testing

How AI Could Replace Animal Testing

New AI technologies show potential to replace animal testing in research, raising ethical and scientific implications. Development is ongoing and unconfirmed for widespread adoption.
future automated home technologies

Ai-Powered Smart Homes: a Glimpse Into Tomorrow’s Living

Smart homes powered by AI are transforming daily living—discover how these innovations will shape your future home and what possibilities await.
ai filters influence attraction

How AI Dating Filters Could Shape Attraction

Ineffective use of AI dating filters may subtly influence perceptions of attractiveness, making you question whether true connection depends on authenticity.