firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Investors and business leaders alike know that diligence and thorough analysis are crucial — but are they enough in the age of AI? A recent live experiment with AI-powered company management suggests that meticulousness alone might not clinch deals or prevent costly mistakes. As AI models are tasked with running a small software firm through its worst week, the results challenge assumptions about what it really takes to succeed in high-pressure environments.

PRIME

Get ready for Prime Big Deal Days — try Prime free

Exclusive member deals on October 6–7, plus fast free delivery. Cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The experiment, conducted by Firmulate, involved four advanced AI models running a simulated small software company facing genuine crises, complex customer interactions, and ethical dilemmas. Each model was tasked with navigating the same scenarios — from customer support challenges to manipulative social engineering attempts — with decisions being carefully versioned and auditable. The goal: determine which AI could best emulate responsible, effective management and close a crucial deal worth €55,000 monthly recurring revenue.

The performance scores reveal a stark reality

  • gpt-5.6-sol led with a perfect score of 95, successfully uncovering critical information buried in company documents and closing the deal.
  • Kimi K3, a newcomer, scored just slightly behind at 93 and was praised for its disciplined approach, refusing manipulation attempts and closing the deal.
  • Sonnet 5 scored 88, managing to close but with some procedural slips and less precision.
  • Fable 5, at 77, also closed the deal but showed more discipline lapses.

Surprisingly, all four models spotted every crisis and refused every manipulation attempt — a critical feat in ensuring integrity. Yet only two models signed the deal based on their analysis. The others, despite correct diagnoses and pitches, left the crucial close on the table.

The hidden weakness lies in the details

Deep within the company’s files, two document references held the key to securing the deal at full price — a fact only the top-performing models read and acted upon. This buried insight made the difference between a successful close and a missed opportunity, emphasizing that reading comprehension and prioritization are vital skills for AI management tools.

Social engineering tests confirm AI integrity

The experiment also tested the models against social engineering tactics, including staged CEO messages and a reporter trick. All five models refused to escalate or approve suspicious requests, with Kimi K3 explicitly reasoning that such requests could be impersonation attempts. This suggests that, at least in controlled conditions, AI can maintain ethical boundaries even under pressure.

The real-world implications

The simulated company, with 13 synthetic employees and real money mechanics, burns €105,000 monthly against a modest €2,300 in monthly revenue. The experiment’s live site — available at firmulate.com/live — demonstrates how AI models behave in real-time business scenarios. This ongoing test bed underscores the critical gap between AI’s diligence and its impact on meaningful business outcomes.

The key lesson: discipline versus impact

The detailed profile of Opus 4.8, the most thorough participant with over 80 learned rules and deep analyses, illustrates that diligence alone doesn’t guarantee success. Despite its exhaustive effort, it left the deal unclosed, mainly due to lapses in escalating issues instead of resolving or documenting them properly. This pattern was visible, albeit weaker, across all models, indicating that volume of rules and analysis isn’t enough — prioritization and disciplined execution matter most.

The broader takeaway for investors

For those investing in AI-driven automation or considering AI tools for their businesses, the experiment offers a vital insight: the ability to read, prioritize, and remain honest under stress outweighs sheer diligence or volume of learned rules. Whether managing customer relationships or navigating crises, AI systems that excel in deep comprehension and disciplined decision-making are more likely to deliver real, measurable impact.

Visit Firmulate’s benchmarks for full results and ongoing updates. Their live platform offers a transparent view of how different models perform in simulated business wargames, helping investors and managers understand what truly matters in AI performance.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

In high-stakes management, diligence must be paired with smart prioritization. AI models that focus deeply, read thoroughly, and stay disciplined under pressure are more likely to win deals and maintain trust — lessons crucial for investors eyeing AI-driven automation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business analysis AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethical dilemma management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Much Printer PPM Speed Do Small Teams Actually Need?

Keen to find out how much printer speed your small team truly needs to stay productive and avoid bottlenecks?

The Orchestration Layer Arrives: What Anthropic’s Finance Agents Mean for Bloomberg, FactSet, and Wall Street

Anthropic releases new agent templates and connectors, positioning Claude as a universal orchestration layer over financial data providers, challenging Bloomberg’s UI dominance.

How Portable Monitor Battery Life Changes Mobile Workflows

Optimizing your portable monitor’s battery life can revolutionize your mobile workflow, but understanding its full impact requires exploring the latest tech advancements.

Can AI Managers Be Trusted? A Live Experiment Reveals Their True Management Personalities

A live experiment tests AI management personalities under pressure, revealing which models can be trusted to finish what they start and make honest decisions in real business crises.