
Get your next haul delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A discount is easy. A trustworthy deal is harder.
For shoppers, the final price matters. For the businesses setting it, a deal also depends on whether an AI agent can spot the right opening, protect customer trust and follow through. Firmulate’s live company experiment puts those abilities under pressure, then asks a practical question for business leaders: how would an AI workforce handle the same challenges inside your company?
One company, one difficult week
In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark also treats a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”
The models recognized every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The gap between identifying a good offer and completing the sale is the story’s sharpest lesson: “Same diagnosis, same pitch — no signature.” For anyone relying on AI to support pricing, sales or customer service, sound advice matters only if an agent can carry it through responsibly.
The clue was already in the company’s files
The deciding competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result points to a familiar business challenge: useful context may exist, but it takes work to find and apply it at the right moment.
Firmulate also tested whether participants would yield to social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its judgment this way: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work still has to end in a decision
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. The league’s note on fairness matters, too: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
The experiment takes place inside a live synthetic company with 13 employees and real money mechanics: €105k in monthly burn against €2.3k MRR, alongside a public cash countdown. Its playbooks contain 680+ self-learned rules, and every workday is versioned. Firmulate says the live company is watchable at firmulate.com. A quiz at firmulate.com uses 242 real, unedited management decisions and invites readers to guess the model.
From watching to a company-specific pilot
A public benchmark can show how models behave in one shared setting. The next step is to see how they handle the pressures and playbooks of a particular business. Firmulate’s enterprise pilot runs crisis scenarios against a digital twin made from a read-only export, then produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.
That offers a way to examine how an AI agent might handle a customer issue, competitor move or manipulation attempt before giving it a role in business operations. For teams considering AI in sales, support or planning, the test is not only whether a model can recommend a deal. It is whether it can find the evidence, respect the rules and make the decision count.

Put your own playbooks to the test
Firmulate’s experiment suggests that recognizing a crisis is only part of the job. A capable AI workforce must also find buried context, maintain trust and follow through. Enterprise teams can run the same wargame against a read-only export of their own business. Explore the pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
