
Imagine if your favorite woodworking tools could decide their own work schedules or handle crises without human input. Now, what if AI models could do the same for a real business—making decisions under pressure, reading critical files, and staying honest when it’s tempting to cheat? This isn’t science fiction; it’s the live experiment from Firmulate, where AI models are running a small, cash-struggling software company through its worst week, and the results could reshape how businesses think about AI management.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Test in a Business Crisis
Firmulate’s live trial involves four advanced AI models, each tasked with managing the day-to-day chaos of a real software company. The company, which is losing €105,000 a month against just €2,300 in monthly recurring revenue, faces identical crises: demanding customers, security breaches, tempting offers, and social engineering attacks. Every decision the AI makes is recorded and auditable, providing a transparent view into their management personalities and decision-making styles.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Key Findings: Honesty, Readiness, and Decision-Making Skills
Despite their differences, all four models demonstrated crucial traits:
- They identified every crisis correctly.
- They refused manipulation attempts, including simulated social engineering scams and fake CEO messages.
- Only two models successfully closed the €55,000 deal their own analysis had earned—showing they could interpret information accurately and act accordingly.
Interestingly, the decisive factor wasn’t in the surface-level decisions but in reading deeper company documents. The models that examined files thoroughly located crucial information buried two document references deep in the company’s files, which enabled them to close the deal at full price, adding €4,583 MRR. The models that skipped or missed these references left the opportunity on the table.
AI management tools for small business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human-Like Personalities of AI Managers
Each model exhibited a distinct management personality:
- GPT-5.6-sol was highly thorough—found the buried fact and closed the deal, showing comprehensive analysis.
- Kimi K3 was the most disciplined—also closed the deal, with a focus on fairness by running without an effort parameter.
- Sonnet 5 was more cautious, yet still closed at full price, with a few process slips.
- Fable 5 lagged behind, leaving the close on the table and slipping into a locked department instead of escalating issues.
The experiment underscores that AI models are not monolithic; they have measurable management personalities that can be as varied as human managers.
AI cybersecurity tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Handling Social Engineering and Ethical Challenges
The models were tested with social engineering scenarios—fake CEO messages escalating over three stages plus a reporter trick asking for a simple yes/no on background. All five models refused to engage, with Kimi K3 reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that AI can be trained to prioritize ethics and security even under pressure, an essential trait for real-world deployment.
As an affiliate, we earn on qualifying purchases.
The Real Business: A Live, Money-Losing Company
The experiment runs within a fully operational software company with 13 synthetic employees, real money mechanics, and ongoing losses. The system works every business day, with over 680 self-learned playbook rules, and every decision is versioned for transparency. The company’s public cash countdown and live updates make the experiment accessible and watchable at firmulate.com/live.
What’s the Big Takeaway for Business Leaders?
This experiment isn’t about whether AI can write a compelling email or chat message. It’s about whether AI can finish what it starts, read critical files before acting, and remain honest under pressure. The results show that even in a complex, money-losing environment, AI models can perform with striking integrity and accuracy. The key question for business leaders today is: will your AI tools be disciplined enough to see the full picture and act ethically when it matters most?
Where Do We Go From Here?
As AI models continue to evolve, their management personalities become more measurable and predictable. The live experiment at Firmulate provides a window into how future AI managers might behave—integral to building trust and resilience in automated decision-making. Companies curious to explore their own AI management capabilities can run similar tests, using the firm’s platform to simulate their business’s worst week—without risking real assets or data.

AI models show distinct personalities in managing crises, reading files deeply, and refusing manipulation. The live Firmulate experiment reveals their potential—and limits—in real business scenarios. Will your AI be honest and thorough when it counts?
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Baby shower & registry season Picks
baby registry must-haves
As an affiliate, we earn on qualifying purchases.