AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When trust is on the line, good intentions need follow-through

Relationships often turn on a small but telling gap: someone says the right thing, understands what is needed, then fails to act. In business, that gap can cost a deal. Firmulate’s live experiment puts AI models in charge of the same small software company during its worst week, testing not only what they notice but whether they follow through under pressure.

Same crises, same company, different decisions

Each model faced the same customers, crises and temptations. Decisions were versioned and auditable. The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. A breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

There was a striking point of agreement: every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The diagnosis and pitch were there; for the others, the signature was not. That difference matters when AI is asked to help run customer relationships, sales or operations. Recognizing the right move is not the same as making it.

The clue was already in the company’s files

The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The experiment’s lesson is practical: a useful AI evaluation has to test whether a system can connect what is happening now to relevant information the company already holds.

Pressure tests revealed more than polished answers

The social-engineering attempts escalated from fake CEO messages over three stages to a reporter asking, “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and tried writing into a locked department instead of escalating. A weaker version of the same discipline problem appeared in all four. Thorough analysis, by itself, did not guarantee a completed job.

The comparison has a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each call.

From watching to a company-specific pilot

The live company has 13 synthetic employees and real money mechanics: it burns €105k/month against €2.3k MRR, shows a public cash countdown, and has accumulated 680+ self-learned playbook rules. Every workday is versioned. The experiment is watchable at firmulate.com.

For an enterprise, the next step is to test its own conditions. Firmulate says a pilot can use a read-only export of a company’s business to run crisis scenarios and produce a board report with a model ranking and the weak points in its playbooks. The pilot writes nothing back to real systems. That offers a way to examine how models might handle a company’s own pressures before trusting them with live work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks under pressure

The experiment suggests that spotting a crisis and refusing a trick are only part of dependable performance. Models also need to use company information, follow through on earned opportunities and escalate when blocked. Enterprises can run the wargame against a read-only export of their own business, with nothing written back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
ai filters influence attraction

How AI Dating Filters Could Shape Attraction

Ineffective use of AI dating filters may subtly influence perceptions of attractiveness, making you question whether true connection depends on authenticity.
enhancing or replacing creativity

AI Writing Assistants: Boosting Creativity or Replacing It?

Just how much can AI writing assistants enhance creativity without replacing it? Discover the balance and implications.
ai therapy insights and gaps

AI in Therapy: What We Know and What We Don’t

Feeling curious about AI’s role in therapy? Discover what we know and what remains uncertain about its impact.
ai enhanced personalized education

AI in Education: Personalized Learning for the Digital Age

Opportunity abounds as AI personalizes education, transforming learning experiences—discover how this innovation can shape your academic future.