
Imagine an AI that’s so thorough it learns over 80 rules, analyzes every angle, and still misses the opportunity to close a deal. Sound familiar? In the world of AI-driven decision-making, diligence alone isn’t enough. Recent experiments with firmulate.com reveal a surprising truth: focus and prioritization often beat volume and exhaustive checks.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test in a Simulated Business Crisis
At the heart of this evaluation was a live, watchable experiment conducted by Firmulate, a platform that simulates real companies facing real crises. Four frontier AI models were tasked with managing a small software company’s worst week—dealing with customer issues, crises, and ethical dilemmas—under identical conditions. Every decision was timestamped, versioned, and auditable, ensuring transparency in how each AI responded.
The Models Compared
- gpt-5.6-sol: The top scorer with a 95/100, successfully identified critical information buried deep in company files and closed the deal.
- Kimi K3: A newcomer with a 93, demonstrating the cleanest discipline and also securing the deal.
- Sonnet 5: Achieved an 88 score, closed the deal but with some process slips.
- Fable 5: With a 77, closed the deal but showed more discipline lapses.
- Opus 4.8: The most thorough participant, learned over 80 rules, yet scored the lowest at 73, ultimately failing to close the deal.
In this setup, all models recognized crises and refused manipulative tactics—like fake CEO messages or staged reporter tricks—showing strong compliance and ethical decision-making. But here’s the twist: only two models actually signed the €55,000 deal, despite all identifying the same opportunity.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Deep in Files Matters
While surface-level analysis seemed enough for most, the key to winning the deal lay hidden two document references deep in the company’s files. The models that accessed and understood this buried information secured the full revenue, underscoring that thoroughness is not just about volume but about reading strategically.
The Social Engineering Test
To test integrity, the models faced staged social engineering attacks—fake CEO messages escalating over multiple stages, and even a reporter’s background request. All five models refused to be manipulated, with Kimi K3 explicitly treating these as impersonation risks. This demonstrates that AI models can maintain ethical boundaries even under pressure.
As an affiliate, we earn on qualifying purchases.
Why Diligence Isn’t Enough: The Lesson for Business AI
Despite Opus 4.8’s ultra-diligent approach—integrating more than 80 learned rules and performing deep analyses—it still finished last in the deal-closure game. The reason? Discipline slipped during the critical close phase, where the team failed to escalate or follow through properly. This highlights a vital lesson for deploying AI in real-world settings: volume of rules and thoroughness do not guarantee impact or success.
Prioritization Over Volume
Across all models, a pattern emerged: those that prioritized reading relevant information and maintaining disciplined decision-making succeeded more than those that simply checked every box or learned more rules. The most thorough model, Opus 4.8, was hampered by a lapse in focus—an important reminder that in business, less can be more if it’s better targeted.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Adoption
For enterprises considering AI tools to handle customer support, sales, or decision-making, the takeaway is clear:
- Does the AI finish what it starts?
- Can it read and understand your critical documents, not just surface data?
- Will it stay honest under pressure?
- And finally, what is the actual cost of useful work—beyond just how much it checks or learns?
The experiment shows that AI models can meet or beat human standards in crisis recognition and ethical integrity. But closing the deal requires disciplined focus, strategic reading, and prioritization—less volume, more impact.
As an affiliate, we earn on qualifying purchases.
Watch the Live Experiment
Curious to see these models in action? Firmulate offers a live, watchable environment where you can observe AI in a simulated business setting, testing decision quality before you hire or deploy it in your own enterprise. Visit firmulate.com/live to explore the ongoing experiments.
The Bottom Line
Effective AI isn’t just about thoroughness; it’s about strategic focus, understanding deeply buried information, and disciplined execution. As this experiment demonstrates, even the most diligent AI can fall short if it overlooks the importance of prioritization. For business leaders, it’s a call to evaluate not just how much AI reads, but how well it reads—and acts.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.