firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your next haul delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A discount is easy. A trustworthy deal is harder.

For shoppers, the final price matters. For the businesses setting it, a deal also depends on whether an AI agent can spot the right opening, protect customer trust and follow through. Firmulate’s live company experiment puts those abilities under pressure, then asks a practical question for business leaders: how would an AI workforce handle the same challenges inside your company?

One company, one difficult week

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Decisions were versioned and auditable. The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The benchmark also treats a breach of trust as decisive: “no amount of good work outweighs a breach of trust.”

The models recognized every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The gap between identifying a good offer and completing the sale is the story’s sharpest lesson: “Same diagnosis, same pitch — no signature.” For anyone relying on AI to support pricing, sales or customer service, sound advice matters only if an agent can carry it through responsibly.

The clue was already in the company’s files

The deciding competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result points to a familiar business challenge: useful context may exist, but it takes work to find and apply it at the right moment.

Firmulate also tested whether participants would yield to social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its judgment this way: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work still has to end in a decision

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. The league’s note on fairness matters, too: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

The experiment takes place inside a live synthetic company with 13 employees and real money mechanics: €105k in monthly burn against €2.3k MRR, alongside a public cash countdown. Its playbooks contain 680+ self-learned rules, and every workday is versioned. Firmulate says the live company is watchable at firmulate.com. A quiz at firmulate.com uses 242 real, unedited management decisions and invites readers to guess the model.

From watching to a company-specific pilot

A public benchmark can show how models behave in one shared setting. The next step is to see how they handle the pressures and playbooks of a particular business. Firmulate’s enterprise pilot runs crisis scenarios against a digital twin made from a read-only export, then produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

That offers a way to examine how an AI agent might handle a customer issue, competitor move or manipulation attempt before giving it a role in business operations. For teams considering AI in sales, support or planning, the test is not only whether a model can recommend a deal. It is whether it can find the evidence, respect the rules and make the decision count.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Firmulate’s experiment suggests that recognizing a crisis is only part of the job. A capable AI workforce must also find buried context, maintain trust and follow through. Enterprise teams can run the same wargame against a read-only export of their own business. Explore the pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Amd Stock

AMD stock surged after the company’s latest earnings report exceeded analyst expectations, signaling positive investor sentiment.

OpenAI in talks to give Trump administration a 5% stake in the company, FT reports

OpenAI is reportedly negotiating to give the Trump administration a 5% ownership stake, according to the Financial Times. Details remain uncertain.

Will Elon Musk Post 40-64 Tweets From September 12 To September 14, 2026?

Speculation surrounds Elon Musk’s potential to post 40-64 tweets from September 12 to 14, 2026, amid rising interest and market signals.

Micron earnings are moments away. Here’s everything investors need to know.

Micron’s quarterly earnings are imminent. Here’s what investors need to know about the upcoming report, confirmed details, and implications.