firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an AI assistant that, despite being the latest in tech, scores just 26 out of 100 in a rigorous test of workplace honesty and reliability. If you run a business, that number should catch your eye. It’s not about how well the AI chats—it’s about whether it can complete tasks, avoid shortcuts, and stay trustworthy when the pressure’s on. A recent public experiment from Firmulate reveals exactly what that looks like—and why even the most advanced models have a way to go before earning your trust.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The Stakes of AI in Business Operations

For owners of small workshops and DIY enthusiasts, the question isn’t just whether AI can generate pretty pictures or write convincing descriptions. It’s whether these models can handle the real-world complexities of running a business—managing crises, making honest decisions, and sticking to protocols. To shed light on this, Firmulate conducted a groundbreaking live experiment, putting four frontier AI models through the same simulated week of a small but busy software company.

Amazon

AI business management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Benchmark Works

The models were tasked with managing a company facing multiple crises—from customer complaints to potential manipulations. Every decision was recorded, versioned, and auditable, ensuring transparency. The goal: see if the AI could identify key issues, resist cheating temptations, and close a profitable deal. The models were tested on their ability to read critical internal documents, recognize hidden facts, and uphold integrity under pressure.

Amazon

AI internal document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings from the Experiment

  • All four models successfully detected every crisis and refused manipulative requests, showing a baseline of cautiousness and compliance.
  • Only two managed to close the deal and sign a €55,000 contract, which they had earned through proper diagnosis and pitch.
  • Interestingly, the decisive advantage lay not in customer interactions, but in the models’ ability to read and understand internal company documents—something that isn’t visible in typical chat demos.
Amazon

AI trustworthiness benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Significance of a 26-Point Baseline

The ‘do-nothing’ baseline score in this test was 26. That might seem low, but it’s a crucial figure. It reflects that even the most minimal effort—like reading internal files or avoiding outright fraud—can earn some points. It underscores that partial progress counts, and the overall score isn’t just about what the AI can do if it’s motivated; it measures whether it will do the right thing even when it’s easiest to cheat.

Amazon

AI compliance and ethics software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Limits of Trust and Performance

One of the core principles in this benchmark is that a single breach of trust caps the total score. This means that if an AI model attempts to manipulate or bypass protocols once, its overall score can’t improve regardless of other successes. It’s an acknowledgment that in real-world business, even a single slip can undermine everything.

What This Means for Small Business Owners

If you’re considering AI to manage customer relations, support, or internal processes, the takeaway is clear: the question isn’t just how well it writes or sounds. It’s whether it can finish what it starts, read the important files, and stay honest under stress. The experiments show that current models are capable of recognizing crises and refusing manipulative tricks—yet they still need improvement to reliably close deals or handle complex internal data.

Behind the Scenes: The Hidden Weakness

Deep in the company’s files, the models that managed to win the deal found critical information that was buried two documents deep. This small but pivotal detail made the difference. It highlights a vital point: the strength of an AI isn’t just in its surface-level responses but in its ability to dig deeper and grasp hidden facts—something essential for trustworthy decision-making in any business.

The Social Engineering Test

The models faced staged manipulations, including fake messages from a CEO and a reporter’s subtle questioning. Every model refused to be manipulated, demonstrating resilience against social engineering tricks—an encouraging sign for businesses wary of fraud and impersonation.

What’s Next and Why It Matters

While the live experiment is ongoing and watchable at firmulate.com/live, the message is already clear: AI models today show promise but still require scrutiny before trusting them with critical business decisions. For small-scale operators and DIYers, it reinforces the importance of testing AI tools thoroughly—using real crises and conflicts—before making them part of your workflow.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Across Two Exhibitions, Jaume Plensa’s Monumental Sculptures Unite Scale And Material

Two concurrent exhibitions showcase Jaume Plensa’s large-scale sculptures, highlighting his mastery of scale and material in public art.

Kennedy Center Surges In Global Coverage

Recent spike in international media coverage highlights rising global interest in the Kennedy Center, with 34 mentions in a recent window, indicating increased attention.

Kennedy Center, Ohio, United States Surges In Global Coverage

The Kennedy Center in Ohio experiences a significant spike in international media mentions, prompting widespread attention and analysis of its current status.

Step Inside The Studios Of 20 Contemporary Black Artists In ‘Another View’

A new exhibition offers an intimate look inside the studios of 20 contemporary Black artists, highlighting their creative processes and cultural impact.