firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine choosing an AI assistant for your business, but instead of promises in demos, you see how it performs amid real crises, file hunts, and ethical tests. The stakes? The difference between a deal closing or losing revenue, and the answer might surprise you. Recently, a live experiment pitted four leading AI models against each other in running a small software company through its toughest week — and the results are shaping the future of AI decision-making.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your next haul delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Live Business War Game: Testing AI Under Real Pressure

In an unprecedented live experiment, four frontier AI models were tasked with managing a small, real-world software company facing its worst week. Every decision, crisis, and temptation was identical for each model, ensuring a fair comparison. This setup isn’t just about chat quality or superficial performance; it’s about how effectively these AIs can handle actual business challenges — reading critical files, resisting manipulations, and closing deals.

Amazon

AI business decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Who Came Out on Top? The Numbers Tell the Story

  • gpt-5.6-sol scored the highest at 95, successfully finding a buried critical fact that led to closing a €55,000 deal, bringing in +€4,583 MRR.
  • Moonshot’s Kimi K3 scored just behind at 93, also closing the deal and exhibiting the cleanest discipline across all models.
  • Sonnet 5 followed at 88, managing to close but with some process slips.
  • Fable 5 and Opus 4.8 scored 77 and 73 respectively, both closing deals but with noticeable lapses.

Notably, all models identified every crisis and refused manipulative attempts — a crucial factor for enterprise trust. The only difference was that the two successful deal-makers read deeper into the company’s files, revealing the hidden information that sealed the deal. This buried fact was two document references deep, stressing the importance of thorough information access.

Amazon

enterprise AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Ethical Vigilance Amid Social Engineering Attacks

During the test, fake CEO messages and reporter tricks were introduced to see if the models would be manipulated. All five models refused to engage, demonstrating strong resistance to social engineering. Kimi K3’s rationale was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that AI’s ethical response is as vital as its analytical skill.

Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real Business, Real Money, Real Risks

The live company managed by these models involved 13 synthetic employees, with real financial mechanics running on a burn rate of €105k per month against €2.3k MRR. The operation is publicly viewable at firmulate.com/live. It is a vivid demonstration of how AI decisions impact ongoing business operations, with every workday’s actions recorded and versioned—an open laboratory of AI management.

Amazon

AI deal closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Role of Deep Document Analysis

The experiment revealed that the decisive weakness of the lower-scoring models was their superficial document reading. The models that delved into the company’s files, uncovering critical buried facts, managed to seal the deal at full price. Conversely, even the most thorough participant, Opus 4.8, left the close on the table partly due to discipline slipping into locked departments instead of escalating issues properly.

Understanding Fairness in AI Performance

It’s important to note that Kimi K3 was run without an effort parameter (the API default), while the other models operated at xhigh effort. This fairness consideration underscores that even with equal effort settings, the models varied significantly in their effectiveness.

Why This Matters for Your Business

Choosing an AI model isn’t just about how well it chats or generates content; it’s about whether it can finish tasks reliably, read critical documents thoroughly, and stay honest under pressure. As AI integrates more deeply into CRM, support, and forecasting, understanding these aspects becomes essential. The public experiment at firmulate.com/benchmarks.html provides a clear, real-world benchmark to guide those choices.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The live AI management experiment shows that the league is open, with the newcomer Kimi K3 just edging out established models in closing deals and maintaining discipline. For businesses, the lesson is clear: pick your AI partner carefully, focusing on reliability, honesty, and the ability to read deeply into your own data — not just on shiny demos or superficial scores.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Der Biomimetische EC-Lüfter Von LONGWELL Erreicht Einen Statischen Wirkungsgrad Von 73-82 % Bei Einer Geräuschreduzierung Von 4-6 dB(A)

LONGWELL’s biomimetic EC fan reaches a static efficiency of 73-82%, reducing noise by 4-6 dB(A), marking a significant advancement in energy-efficient ventilation technology.

Hedgehog USA Signs Strategic Energy Supply Agreement With Idemitsu To Advance Powered Land Data Center Platform In Texas

Hedgehog USA partners with Idemitsu in a strategic energy supply agreement to support its land data center platform in Texas, advancing its power infrastructure.

Meta’s Consumer-Focused AI Agent Could Be Weeks From Launch

Meta is preparing to release a new AI-powered consumer assistant within weeks, marking a significant step in its AI strategy.

Neuromorphic Chips: Mimicking the Human Brain in Silicon

Learn how neuromorphic chips mimic the human brain in silicon, revolutionizing AI and smart devices—discover what makes them so groundbreaking.