
Get sports gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
What Sports Can Teach Us About AI’s True Performance
Imagine a coach who, despite perfect strategy and flawless execution, can’t score above 26 points in a crucial game — even when playing against an empty field. This isn’t a sports fantasy, but a real story in AI management testing. In a recent public experiment by Firmulate, AI models faced a simulated week of running a small software company. The results reveal not only how AI handles crises but also expose the limits of current benchmarks and what you should really look for before trusting AI with your business.
As an affiliate, we earn on qualifying purchases.
The Reality Behind AI Benchmarks
In the world of artificial intelligence, scores matter — but only if they reflect real-world performance. Recently, a public experiment known as the Crucible League tested four top AI models by running them through the worst week a small company could face: customer crises, internal manipulations, and complex decision-making. Every model was given the same scenario, with identical customers, crises, and temptations. The goal? See if they could handle real management challenges, not just generate convincing chat.
The Surprising Baseline Score
One key finding was that even a do-nothing baseline — a model that makes no effort — scores 26 points out of a possible 100. This partial score might seem low, but it’s a crucial benchmark. It shows that simply doing nothing is better than outright failure, and that partial progress counts. More importantly, a single breach of trust, like accepting a manipulation or ignoring critical files, caps the total score at 26, regardless of other successes. This honesty in scoring prevents inflated ratings and ensures benchmarks reflect true management integrity.
What the Experiment Revealed
All four AI models identified every crisis — from customer complaints to internal manipulations. They refused every manipulation attempt, including social engineering tricks like fake CEO messages or reporter traps. For example, five out of five models refused to approve a fake request presented as a background check, citing suspicion of impersonation. This indicates that AI models are becoming adept at recognizing deceptive tactics and maintaining ethical boundaries.
The Real Win: Reading the Files
However, the true differentiator was not just crisis management but how deeply each model could analyze company documents. The decisive weakness was buried two references deep within the company’s files, not in the immediate customer interactions. Models that read and understand these internal files managed to close a €55,000 deal at full price — adding an extra €4,583 MRR to the company’s revenue. Conversely, some models, despite diagnosing and pitching well, left the deal on the table due to a failure to read deeper information.
The Limits of Current Benchmarks
What does this mean for business managers? It shows that current AI benchmarks often focus on superficial skills like generating convincing language but neglect critical capabilities such as deep document comprehension and trustworthiness under pressure. The experiment also included a live, watchable simulation where 13 synthetic employees, powered by AI, managed real money — burning €105,000 monthly against a revenue of just €2,300. These models learned over 680 self-imposed rules and were continually versioned, simulating real management decisions in a complex environment.
The Importance of Discipline and Trust
In the experiment, even the most thorough participant, Opus 4.8, which analyzed over 80 rules, left significant opportunities unseized. It failed to escalate when discipline slipped, leaving a close opportunity on the table. This highlights that discipline and accountability are as critical as intelligence in AI management — qualities that current models are still developing.
What Should Business Leaders Take Away?
When evaluating AI for management tasks, the key questions aren’t just about how well it writes or chats. Instead, focus on whether it completes what it starts, reads the necessary internal data, and stays honest when under pressure. The experiment underscores that a high score isn’t just about superficial competence; it’s about trustworthiness, depth of understanding, and discipline.

Key Takeaways for Business Leaders
- AI benchmarks often overstate capabilities; real-world performance depends on trust, depth, and discipline.
- A score of 26 from a do-nothing baseline highlights the importance of partial progress and safeguards against overconfidence.
- Deep internal understanding, not just crisis recognition, separates truly effective AI managers from the rest.
- Trustworthiness and adherence to rules are vital — even in AI, breaches cap potential gains.
- Before deploying AI in critical functions, consider running live wargames like Firmulate’s to test actual management skills.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
