
Imagine a company with no human employees, running its operations every single day, yet losing €105,000 each month against a tiny €2,300 monthly recurring revenue. Now, imagine watching that company’s decisions unfold in real-time, with every move scrutinized and challenged. This is not a fictional story but a groundbreaking live experiment revealing how AI models behave under pressure — and what it means for your digital future.
The Live Company in Action
At the heart of this experiment is Firmulate, a real, functioning software company run entirely by artificial intelligence models. Every weekday, 13 synthetic employees make decisions, respond to crises, and attempt to close deals—all in a public, transparent environment that resets and archives each day’s decisions.
What makes this setup extraordinary is its transparency and realism. The company faces the same customer demands, crises, and temptations as any real business. It burns through €105,000 each month, chasing a mere €2,300 in monthly revenue, with a public cash countdown ticking down daily. Every decision is version-controlled and auditable, allowing observers to track exactly how each AI model responds under pressure.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Self-Learning, Self-Reporting
Four frontier AI models, including the well-known GPT-5.6, Kimi K3, Sonnet 5, and Opus 4.8, were each tasked with running this business through its worst week. They faced the same customer issues, internal crises, and manipulative tactics—yet their responses varied significantly.
All models successfully identified and responded to crises, and all refused to engage in manipulation attempts such as fake CEO messages or secret approvals. However, only two managed to close a deal at full price, earning an additional €4,583 in monthly recurring revenue, demonstrating the importance of thorough information reading and disciplined decision-making.

AI Automation Mastery: Learn AI Automation, Build Smart Systems, Master No-Code & AI Tools, Boost Productivity, and Turn Your Skills into Income
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Numbers Say
- Model Scores: GPT-5.6 scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, and Opus 4.8 scored 77.
- Deal Closure: Only GPT-5.6 and Kimi K3 managed to close the €55,000 deal they identified during their analysis.
- Key Weakness: The decisive factor was the models’ ability to uncover critical information buried two document references deep in the company’s files—an insight that won the full-price deal.
- Human-Like Deception Resistance: All models refused to be duped by staged social engineering attacks, including fake CEO messages and reporter tricks.

The AI Documentation Ethics Audit Kit: A 7-Question Framework for Grading, Fixing, and Future-Proofing Your AI Product Documentation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Business
As AI begins to integrate more deeply into business operations—handling customer support, decision-making, and process automation—the question isn’t merely whether it can generate convincing chat responses. The real measure is whether it can finish what it starts, stay honest under pressure, and act based on a comprehensive understanding of internal data.
The live experiment shows that high-performing AI models can resist deception and identify hidden information when given access to the right documents. But even the best still leave opportunities on the table—like the Opus 4.8, which had the most thorough analysis but still slipped in closing the deal due to process slips and discipline lapses.

Games of Deception: The True Story of the First U.S. Olympic Basketball Team at the 1936 Olympics inHitler's Germany
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Trust and Discipline Make or Break AI Success
Trustworthiness isn’t just about honesty; it’s about consistent discipline and thoroughness. The experiment’s results highlight that models with more comprehensive learned rules—like Opus—are better equipped to handle crises but still need disciplined processes to convert insights into action. Interestingly, the models that ran without effort parameters, like Kimi K3, performed the cleanest, suggesting that simplicity in AI behavior can sometimes yield better discipline.
Watch It Live and Decide
If you’re curious how AI can perform in real business scenarios—and whether it’s ready to be trusted with your operations—you can observe this experiment in real-time at firmulate.com/live. The company runs every weekday, showing decision-making under pressure, crisis management, and the potential pitfalls that come with AI automation.
Additionally, you can explore the detailed results and plain-language insights from the experiment at firmulate.com/quotes.html, or test your management intuition with a quiz at firmulate.com/quiz.html. For enterprise users, there’s also a pilot mode to run your own business scenarios without affecting your real systems.

This live experiment exposes how AI models handle crises, honesty, and decision-making under pressure in a real business context. Trust and discipline are key to harnessing AI’s potential—watch it unfold at firmulate.com/live and see what the future of AI-driven management might look like.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html