firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine choosing an AI assistant for your business, but instead of promises in demos, you see how it performs amid real crises, file hunts, and ethical tests. The stakes? The difference between a deal closing or losing revenue, and the answer might surprise you. Recently, a live experiment pitted four leading AI models against each other in running a small software company through its toughest week — and the results are shaping the future of AI decision-making.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your next haul delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Live Business War Game: Testing AI Under Real Pressure

In an unprecedented live experiment, four frontier AI models were tasked with managing a small, real-world software company facing its worst week. Every decision, crisis, and temptation was identical for each model, ensuring a fair comparison. This setup isn’t just about chat quality or superficial performance; it’s about how effectively these AIs can handle actual business challenges — reading critical files, resisting manipulations, and closing deals.

Amazon

AI business decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Who Came Out on Top? The Numbers Tell the Story

  • gpt-5.6-sol scored the highest at 95, successfully finding a buried critical fact that led to closing a €55,000 deal, bringing in +€4,583 MRR.
  • Moonshot’s Kimi K3 scored just behind at 93, also closing the deal and exhibiting the cleanest discipline across all models.
  • Sonnet 5 followed at 88, managing to close but with some process slips.
  • Fable 5 and Opus 4.8 scored 77 and 73 respectively, both closing deals but with noticeable lapses.

Notably, all models identified every crisis and refused manipulative attempts — a crucial factor for enterprise trust. The only difference was that the two successful deal-makers read deeper into the company’s files, revealing the hidden information that sealed the deal. This buried fact was two document references deep, stressing the importance of thorough information access.

Amazon

enterprise AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Ethical Vigilance Amid Social Engineering Attacks

During the test, fake CEO messages and reporter tricks were introduced to see if the models would be manipulated. All five models refused to engage, demonstrating strong resistance to social engineering. Kimi K3’s rationale was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that AI’s ethical response is as vital as its analytical skill.

Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real Business, Real Money, Real Risks

The live company managed by these models involved 13 synthetic employees, with real financial mechanics running on a burn rate of €105k per month against €2.3k MRR. The operation is publicly viewable at firmulate.com/live. It is a vivid demonstration of how AI decisions impact ongoing business operations, with every workday’s actions recorded and versioned—an open laboratory of AI management.

Amazon

AI deal closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Role of Deep Document Analysis

The experiment revealed that the decisive weakness of the lower-scoring models was their superficial document reading. The models that delved into the company’s files, uncovering critical buried facts, managed to seal the deal at full price. Conversely, even the most thorough participant, Opus 4.8, left the close on the table partly due to discipline slipping into locked departments instead of escalating issues properly.

Understanding Fairness in AI Performance

It’s important to note that Kimi K3 was run without an effort parameter (the API default), while the other models operated at xhigh effort. This fairness consideration underscores that even with equal effort settings, the models varied significantly in their effectiveness.

Why This Matters for Your Business

Choosing an AI model isn’t just about how well it chats or generates content; it’s about whether it can finish tasks reliably, read critical documents thoroughly, and stay honest under pressure. As AI integrates more deeply into CRM, support, and forecasting, understanding these aspects becomes essential. The public experiment at firmulate.com/benchmarks.html provides a clear, real-world benchmark to guide those choices.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The live AI management experiment shows that the league is open, with the newcomer Kimi K3 just edging out established models in closing deals and maintaining discipline. For businesses, the lesson is clear: pick your AI partner carefully, focusing on reliability, honesty, and the ability to read deeply into your own data — not just on shiny demos or superficial scores.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Photonic Chips: Computing With Light to Shatter Speed Limits

Just as light revolutionizes computing, photonic chips promise unprecedented speed and efficiency that could transform technology forever.

Fiserv To Release Second Quarter Earnings Results On August 6, 2026

Fiserv will release its second quarter earnings results on August 6, 2026, providing insights into its financial performance for the period.

Tercera Launches New Advisory Practice, Complementing Existing Investment Arm

Tercera announces the launch of a new advisory division alongside its existing investment business, expanding its services and strategic offerings.

Space Tourism: The Final Frontier for Travelers

Curious about the exhilarating world of space tourism? Discover how you can embark on an unforgettable journey beyond Earth’s atmosphere!