
Imagine choosing an AI assistant for your business, but instead of promises in demos, you see how it performs amid real crises, file hunts, and ethical tests. The stakes? The difference between a deal closing or losing revenue, and the answer might surprise you. Recently, a live experiment pitted four leading AI models against each other in running a small software company through its toughest week — and the results are shaping the future of AI decision-making.
Get your next haul delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Live Business War Game: Testing AI Under Real Pressure
In an unprecedented live experiment, four frontier AI models were tasked with managing a small, real-world software company facing its worst week. Every decision, crisis, and temptation was identical for each model, ensuring a fair comparison. This setup isn’t just about chat quality or superficial performance; it’s about how effectively these AIs can handle actual business challenges — reading critical files, resisting manipulations, and closing deals.
AI business decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Who Came Out on Top? The Numbers Tell the Story
- gpt-5.6-sol scored the highest at 95, successfully finding a buried critical fact that led to closing a €55,000 deal, bringing in +€4,583 MRR.
- Moonshot’s Kimi K3 scored just behind at 93, also closing the deal and exhibiting the cleanest discipline across all models.
- Sonnet 5 followed at 88, managing to close but with some process slips.
- Fable 5 and Opus 4.8 scored 77 and 73 respectively, both closing deals but with noticeable lapses.
Notably, all models identified every crisis and refused manipulative attempts — a crucial factor for enterprise trust. The only difference was that the two successful deal-makers read deeper into the company’s files, revealing the hidden information that sealed the deal. This buried fact was two document references deep, stressing the importance of thorough information access.
enterprise AI document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Ethical Vigilance Amid Social Engineering Attacks
During the test, fake CEO messages and reporter tricks were introduced to see if the models would be manipulated. All five models refused to engage, demonstrating strong resistance to social engineering. Kimi K3’s rationale was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that AI’s ethical response is as vital as its analytical skill.
AI ethical decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real Business, Real Money, Real Risks
The live company managed by these models involved 13 synthetic employees, with real financial mechanics running on a burn rate of €105k per month against €2.3k MRR. The operation is publicly viewable at firmulate.com/live. It is a vivid demonstration of how AI decisions impact ongoing business operations, with every workday’s actions recorded and versioned—an open laboratory of AI management.
As an affiliate, we earn on qualifying purchases.
The Surprising Role of Deep Document Analysis
The experiment revealed that the decisive weakness of the lower-scoring models was their superficial document reading. The models that delved into the company’s files, uncovering critical buried facts, managed to seal the deal at full price. Conversely, even the most thorough participant, Opus 4.8, left the close on the table partly due to discipline slipping into locked departments instead of escalating issues properly.
Understanding Fairness in AI Performance
It’s important to note that Kimi K3 was run without an effort parameter (the API default), while the other models operated at xhigh effort. This fairness consideration underscores that even with equal effort settings, the models varied significantly in their effectiveness.
Why This Matters for Your Business
Choosing an AI model isn’t just about how well it chats or generates content; it’s about whether it can finish tasks reliably, read critical documents thoroughly, and stay honest under pressure. As AI integrates more deeply into CRM, support, and forecasting, understanding these aspects becomes essential. The public experiment at firmulate.com/benchmarks.html provides a clear, real-world benchmark to guide those choices.

The live AI management experiment shows that the league is open, with the newcomer Kimi K3 just edging out established models in closing deals and maintaining discipline. For businesses, the lesson is clear: pick your AI partner carefully, focusing on reliability, honesty, and the ability to read deeply into your own data — not just on shiny demos or superficial scores.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
