
Imagine a sports team that plays every game live, with fans watching every move, yet the team has no players and is losing every single match. This is the reality of a groundbreaking experiment in artificial intelligence, where a company runs entirely on AI models scrutinized for honesty, discipline, and decision-making — and every step is on display for the world to see.
The Live Experiment: A Company Without Employees
At the heart of this experiment is a real, functioning software company, monitored daily at firmulate.com/live. No humans, no shortcuts — just 13 synthetic employees powered by advanced AI models. Every workday, the company faces real crises, customer requests, and ethical dilemmas, all while burning €105,000 a month against a modest €2,300 monthly recurring revenue.
This setup isn’t just a stunt; it’s a rigorous test of AI decision-making under real-world pressures. The models are designed to operate transparently, with every decision versioned and auditable, and rules learned and refined over time — more than 680 self-learned playbook rules in total.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Are These AI Models Trying to Achieve?
The experiment pits four frontier AI models—each with different approaches—against the same tough week of business. They confront customer crises, internal mistakes, and even social engineering attempts designed to manipulate decisions. All models successfully identify crises and refuse manipulative tactics, demonstrating a baseline of honesty and discipline.
However, when it comes to closing deals, differences emerge. For example, only two models signed a €55,000 contract their own analysis recommended. The critical factor? An overlooked detail in the company’s own files — a buried fact that, if read, would have secured the deal at full price (+€4,583 MRR). The model that discovered this won decisively, proving that reading and understanding internal documentation is vital for success.
Social Engineering and Ethical Resilience
The experiment also tested the models’ resistance to social engineering. Fake CEO messages and a staged reporter request were used to see if the models would be fooled into bypassing controls. All five tested models refused to be manipulated, with one explicitly treating suspicious requests as impersonation risks. This resilience suggests that AI models trained with strict protocols can maintain integrity even under pressure.
The Reality of a Money-Losing Business
Despite its advanced decision-making, the company is not profitable — it continues to burn through €105k every month. The live dashboard displays the company’s public cash countdown, emphasizing its fragile financial state. This isn’t a theoretical exercise but a live, ongoing story of survival, with every decision, misstep, and victory open for observation.
The Performance League: Who Comes Out on Top?
- gpt-5.6-sol scored 95, identified the hidden fact, and closed the deal at full price.
- Kimi K3 scored 93, closed the deal, and displayed the cleanest discipline among models.
- Sonnet 5 scored 88, finished the deal but with minor slips.
- Fable 5 scored 77, also closed the deal but less consistently.
All models could recognize crises and refuse manipulation, but only the top performers could turn insights into profitable deals consistently.
Implications for Business and Technology
This experiment highlights a critical point for managers: the difference between AI chat capabilities and actual decision-making. It’s not about how well an AI can write or chat but whether it can finish what it starts, stay honest, and read internal data—skills essential for real-world business success. The experiment proves that AI models can be trained to uphold discipline and integrity, but the challenge remains in translating that into consistent profitability.
Takeaway
The live company experiment is a raw, unfiltered view of AI’s potential and limits. It demonstrates that AI can be honest, diligent, and resistant to manipulation under pressure, but profitability depends on more than decision accuracy. For businesses considering AI automation, the key takeaway is that performance must be judged on results—reading internal data, following through, and maintaining ethical standards—just as in sports, where discipline often outweighs talent alone.

This live experiment offers a rare, transparent look at AI decision-making in a real business, showing how honesty, discipline, and thorough reading are crucial for success — not just chat skills.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html