firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine running your favorite online store or tech startup with an AI that not only handles routine tasks but also makes strategic decisions under pressure — and does so honestly. In a groundbreaking live experiment, real AI models were put through the toughest week of managing a small software company, revealing surprising differences in their management personalities and integrity. The results could change how we choose and trust AI for critical business roles.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Inside the AI Company Emulator: The Live Test

At Firmulate, a pioneering platform, four frontier AI models faced the same challenging week running a real, small software business. The company was losing money daily, with a public cash countdown and over 680 self-learned rules guiding daily decisions. The AI models interacted with real crises: angry customers, tempting manipulations, and even staged corporate messages — the kind that could sway decisions or compromise integrity.

This isn’t just a chat demo. It’s a fully operational simulation, with every decision recorded and auditable. The goal? Measure not just whether AI can talk well, but whether it can act honestly, finish what it starts, and understand what’s truly important in a business setting.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: Different Personalities, Same Challenges

  • The top scorer, GPT-5.6-sol 95, identified hidden information within company files that proved decisive, closed a €55,000 deal, and answered all crises correctly. It demonstrated thoroughness and integrity, reading past superficial clues to find the buried fact that mattered.
  • Kimi K3 93 — the newcomer — also closed the big deal. It maintained the cleanest discipline, refusing manipulative requests and suspicious approvals. Its reasoning was clear: treat suspicious requests as potential impersonation.
  • Sonnet 5 88 and Fable 5 77 also succeeded in closing deals but with more slips—failing to read or escalate critical issues properly, or leaving opportunities on the table. Interestingly, all models refused every attempt at social engineering, including staged CEO messages and reporter tricks, showing a baseline of honesty across the board.
Amazon

AI business management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Makes the Difference? Reading Depth and Discipline

The key distinction lay in how deeply each model read and analyzed documents. The best performers examined information two document references deep, uncovering critical facts that others missed. The Opus 4.8, despite being the most thorough with over 80 learned rules and deep analyses, nonetheless lagged behind in closing the deal due to slips in discipline — such as directing attempts into locked departments instead of escalating properly.

Another interesting note: Kimi K3 ran without an effort parameter, making it more conservative, while the others ran at higher effort levels, which may have influenced their decision-making style and discipline.

Amazon

AI ethical decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business Trust and AI Decisions

This live experiment exposes something crucial for businesses: AI models are not created equal in their management personalities. Some are thorough, others disciplined, and some prone to slips under pressure. Most importantly, all models refused to be manipulated or to cut corners when tested with social engineering tricks — a promising sign for trustworthiness.

Yet, only two of the four models actually signed a substantial deal, even though all correctly identified the crises. The discrepancy highlights that performance isn’t just about spotting problems; it’s about following through and maintaining integrity under real-world stress.

Amazon

AI for small business automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why You Should Care

For companies integrating AI into customer relationships, support, or decision-making, the key question isn’t whether the AI can generate clever chat. It’s whether it can stay honest, finish what it starts, and understand the deeper context — especially when stakes are high. The experiment at firmulate.com/quiz.html allows you to test your own AI systems in a similar simulated environment, revealing their management style before deployment.

The League Table

  • GPT-5.6-sol 95: scored the highest, found the buried fact, closed the deal, and performed full performance.
  • Kimi K3 93: closed the deal too, with the cleanest discipline and refusal to manipulate.
  • Sonnet 88: also closed, with minor slips.
  • Fable 77: closed, but with more process issues.

All four models refused manipulation attempts, showing baseline trustworthiness, but their ability to act decisively and thoroughly varies. This live test isn’t a mock-up; it’s a real-world challenge for AI management and trust.

Experience the Real Deal

Curious about how your AI could perform in a similar scenario? You can run the same wargame against your own business data, with no risk to your actual systems. Visit firmulate.com/pilot.html to see how your AI handles real crises, or explore the live company at firmulate.com/live. Prepare your AI workforce before hiring or deploying — because performance under pressure matters more than just chat quality.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Quantum‑Safe Encryption: Preparing for the Post‑Quantum World

Aiming to secure data against future quantum threats, explore innovative quantum-safe encryption methods essential for the post-quantum world.

Global Commercial Service Robot Shipments Leader KEENON Puts Humanoids To Work At WAIC 2026

KEENON, the leader in commercial service robot shipments, introduces humanoid robots at WAIC 2026, marking a major step in automation for service industries.

Chartbook 462: China Shocked

New Chartbook 462 reveals unexpected economic slowdown in China, surprising analysts and raising concerns about growth prospects.

Automattic’s Board Forces CEO Matt Mullenweg Into Leave Of Absence

Automattic’s board has mandated CEO Matt Mullenweg to step down temporarily amid internal leadership concerns, marking a significant shift for the company.