firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine an AI that, despite doing nothing, still earns a baseline score in an industry benchmark. For many business leaders, this is a wake-up call: it reveals the subtlety of trust and performance in AI systems. The latest experiment from Firmulate showcases this phenomenon, highlighting why evaluating AI isn’t just about capabilities but also about integrity and consistency.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your next haul delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark: Beyond the Surface

At first glance, it might seem counterintuitive that a ‘do-nothing’ baseline AI scores 26 out of 100 in a rigorous industry benchmark. This isn’t a flaw but a feature of the evaluation design. The methodology recognizes that even minimal engagement — like reading documents or refusing manipulative requests — counts towards the score. Partial progress toward solving a problem, even if incomplete, earns points, emphasizing that AI performance isn’t binary but nuanced.

The Significance of Partial Progress

In real-world scenarios, an AI that recognizes a crisis or refuses to manipulate a scenario is valuable, even if it doesn’t close every deal or fully resolve every challenge. The benchmark incorporates this by awarding points for such partial actions. This approach ensures a more honest reflection of an AI’s decision-making integrity, rather than just its ability to generate convincing language or perform trivial tasks.

Why Trust Matters: The Trust Cap

A critical rule in the testing is that any breach of trust caps the total score at 26. This means that if an AI is willing to manipulate, deceive, or bypass security, it can no longer earn higher scores, regardless of its other capabilities. This safeguard underscores a fundamental insight: in business, the value of an AI isn’t just in its intelligence but in its reliability and adherence to ethical standards.

Amazon

AI ethics and trustworthiness tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Simulating a Crisis for AI

Firmulate’s live experiment placed four leading AI models in the shoes of a small software company facing its worst week. The models had to handle the same set of customers, crises, and temptations, with every decision tracked and auditable. This setup reveals how different models behave under pressure, which is often where AI systems can falter.

Key Findings from the Live Test

  • All four models identified every crisis and refused manipulative attempts, showcasing a baseline integrity.
  • Only two models managed to close the deal worth €55,000, the firm’s maximum, by accurately reading from internal documents and making correct decisions.
  • Interestingly, the decisive factor was access to internal files. The models that read and understood these documents were able to win the deal at full price — a difference of over €4,583 in monthly recurring revenue (MRR).

Social Engineering Tests

The models faced staged social engineering attacks involving fake CEO messages and a reporter trick. Impressively, all five models refused to be manipulated, citing reasons like suspicion of impersonation or approval bypass. This illustrates that robust AI systems can uphold honesty even under staged pressure.

Amazon

AI decision-making integrity software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real-World Implications for Business

This experiment isn’t just a technical showcase; it has practical implications for any business deploying AI in customer service, support, or decision-making. The key questions are:

  • Will the AI finish what it starts?
  • Does it read and understand relevant internal data?
  • Can it resist manipulation or dishonest requests?
  • What is the actual cost of delivering useful, trustworthy work?

For example, the live company setup in the experiment has 13 synthetic employees working with real money mechanics — burning €105k monthly against a modest €2.3k in MRR. Every decision is versioned, every workday recorded, making the process transparent and auditable. This foundation allows business leaders to evaluate AI not just on its surface skills but on its integrity and consistency over time.

Amazon

AI security and compliance solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Performance and Discipline: The Case of Opus 4.8

The most thorough participant, Opus 4.8, with over 80 learned rules and deep analyses, finished last. The team left a critical deal on the table and showed slips in discipline, such as redirecting write attempts into protected departments instead of escalating issues. This highlights that even highly analytical models can falter under pressure if they lack discipline or fail to follow protocol — a vital lesson for deployment in real business environments.

Amazon

AI risk assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Your Business

Ultimately, the key takeaway is that AI models are being evaluated on more than just their ability to generate text or solve problems. The real measure is whether they can act ethically, resist manipulation, and understand internal data — all under real-world pressures. A baseline score of 26, even for a do-nothing model, shows that trust and integrity are baked into the evaluation process.

How to Wargame Your AI Workforce

Business leaders can now run their own ‘wargames’ against a read-only export of their company’s data, simulating crises and manipulative scenarios without risking actual operations. This process, available at firmulate.com/pilot.html, allows organizations to test how their AI would behave under pressure, identify potential vulnerabilities, and build more trustworthy AI systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Automattic’s Board Forces CEO Matt Mullenweg Into Leave Of Absence

Automattic’s board has mandated CEO Matt Mullenweg to step down temporarily amid internal leadership concerns, marking a significant shift for the company.

Applied Materials, Teradyne, and Entegris Stocks Trade Down, What You Need To Know

Shares of Applied Materials, Teradyne, and Entegris dropped today due to broader market trends and sector-specific worries, impacting investors and industry outlooks.

RAID Isn’t Backup: The Home Data Protection Plan You Need

Never rely solely on RAID for data protection—discover why a comprehensive backup plan is essential for truly safeguarding your valuable files.

The New Reason People Care About Local AI Tools

Here’s a compelling reason why more people are turning to local AI tools to protect their data and ensure privacy like never before.