AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

In sport, a team can read the play perfectly and still fail to execute it. That gap between seeing the opening and taking it is now showing up in AI-run businesses, too. Firmulate’s latest company wargame put frontier models through the same rough week. Every model spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get sports gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The lesson for anyone bringing AI into a club, venue or sports business is practical: a convincing answer is not the same as a sound decision under pressure. Firmulate makes that difference visible in a live, watchable experiment.

The same match, different decisions

In the final Crucible League, published in July 2026, gpt-5.6-sol finished first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The models faced the same small software company, customers, crises and temptations. Their decisions were versioned and auditable. The test’s sharpest contrast was not crisis recognition: all models identified every crisis. It was follow-through. Only two completed the €55,000 deal that their own analysis had justified. Same diagnosis, same pitch — no signature.

The clue was already in the files

The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The finding is a reminder that useful business judgment can depend on connecting evidence that is present but easy to overlook.

Integrity faced its own test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

A strong analyst can still leave the play unfinished

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. It left the deal unsigned and discipline slipped: it tried to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four models.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The ranking is the published result, but that difference belongs alongside it when readers judge the contest.

A company you can watch

The experiment sits within a live synthetic company with 13 employees and real money mechanics: €105,000 in monthly burn against €2,300 in monthly recurring revenue, alongside a public cash countdown. Its playbook has grown to more than 680 self-learned rules, and every workday is versioned. Firmulate also offers a quiz built from 242 real, unedited management decisions, asking visitors to guess which model made each call.

For sports organizations, the point is not that a synthetic software company predicts how an AI will manage a team. It offers a way to examine the kinds of decisions an AI workforce might face: spotting a crisis, finding relevant information, respecting boundaries and carrying a justified plan through to completion.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to trying it yourself

Firmulate’s enterprise pilot applies the wargame to a read-only export of a company’s own business. It puts crisis scenarios against that company context and produces a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Wave Height Calibration: Engineering Secrets Revealed

Bridging the gap between inaccurate wave measurements and reliable data, discover engineering secrets that can transform your calibration process.

AI Management Skills Show Their True Color in Live Business Simulation

A live experiment shows AI models managing a real company’s worst week, revealing that management skills like resilience and honesty are beyond what chat benchmarks measure.

Sustainable Heat Recovery Systems in Surf Lagoons

Jump into sustainable heat recovery systems in surf lagoons to learn how they can transform energy efficiency and reduce environmental impact.

Scroll-Linked SVG Scenes: A Look Inside “The Paper Fox — a bedtime story in seven scenes” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“The Paper…