AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A stress test for more than batteries

Home-energy readers understand the value of testing a system before the difficult day arrives. Backup power should be ready before an outage; an AI agent trusted with customers, contracts or company data should face its own pressure test before a real emergency exposes its weaknesses.

Firmulate has produced an unusually encouraging result. Five frontier AI models encountered escalating messages from a fake chief executive and a reporter seeking confidential confirmation. All five refused every manipulation attempt. The episode suggests that integrity under pressure is something enterprises can examine before production—not merely reconstruct later in an incident report.

Preventing Cheating Through Academic Integrity (Quick Reference Guide)

Preventing Cheating Through Academic Integrity (Quick Reference Guide)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same bad week for every model

Firmulate gave each model control of the same small software company during its worst week. The customers, crises and temptations remained constant, while every decision was versioned and auditable. The synthetic company has 13 employees and deliberately uncomfortable finances: burn of €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown adding urgency.

That urgency made the social-engineering challenge more realistic. The fake CEO messages escalated over three stages, pressing the models to bypass normal safeguards. A separate reporter tried a softer route, asking for “just one yes/no, on background.” The result did not vary: 5 of 5 models refused.

Kimi K3 stated the danger plainly in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because the request was framed as executive authority combined with time pressure—the combination that can tempt a helpful system to treat compliance as success. More examples of the models’ recorded language appear on Firmulate’s public quotes page.

Refusal was necessary, but it was not the whole job

The broader experiment exposed a useful distinction between staying safe and completing valuable work. Every model spotted every crisis and resisted every manipulation attempt, yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap as: “Same diagnosis, same pitch — no signature.”

The commercial opportunity depended on a detail hidden two document references deep in the company’s own files rather than in the customer event. Models that found it won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson is not simply that an agent should refuse suspicious instructions. It must also read the available evidence, pursue legitimate work and finish what it starts.

The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust.”

K3’s performance carries an important fairness note. It ran with the API default and without an effort parameter, while the other participants ran at xhigh. That difference should remain visible when readers compare the results.

Thoroughness did not guarantee first place

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, though less strongly.

This is why polished answers alone are a poor readiness test. A model can identify the danger, explain the right move and still fail to carry the legitimate task across the finish line. Conversely, a productive agent is not acceptable if it yields to impersonation or pressure. Firms need both integrity and execution.

The live company continues with real money mechanics, a public cash countdown and more than 680 self-learned playbook rules. Its workdays remain versioned, making the experiment watchable rather than retrospective. Firmulate also uses 242 real, unedited management decisions in a quiz that asks people to guess which model made each choice.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
THE AI GOVERNANCE ARCHITECT: BUILDING MODEL RISK MANAGEMENT AND COMPLIANCE FRAMEWORKS: A Practitioner's Blueprint for Auditable MLOps, Systemic Traceability, and Scaling Trust in Regulated Enterprise

THE AI GOVERNANCE ARCHITECT: BUILDING MODEL RISK MANAGEMENT AND COMPLIANCE FRAMEWORKS: A Practitioner's Blueprint for Auditable MLOps, Systemic Traceability, and Scaling Trust in Regulated Enterprise

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the human pressure points before deployment

For businesses adopting AI agents—and for energy companies considering them across customer service, sales or forecasting—the reassuring result is that all five models resisted the fake executive and the reporter. The caution is that refusal alone does not establish readiness.

A credible evaluation should combine manipulation attempts with ordinary commercial work: Can the agent protect trust, consult the company’s own records, escalate when permissions block it and complete a justified action? Firmulate’s experiment shows that these behaviors can be observed under controlled pressure. The best time to discover whether an AI obeys a fake boss—or abandons a legitimate deal—is before either situation becomes real.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI-Powered Safety: Streamlined EHS Operations for Managers

AI-Powered Safety: Streamlined EHS Operations for Managers

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Modern AI Agent with Claude AI: A Practical Guide to Building Autonomous Workflows for Real-World Use

The Modern AI Agent with Claude AI: A Practical Guide to Building Autonomous Workflows for Real-World Use

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Grounding Questions for Battery Backup: What You Actually Need to Know

Learn the essential grounding considerations for battery backups and discover what you actually need to know to ensure safety and compliance.

Splitters, Adapters, and “Cheater Plugs”: What’s Safe in an Outage?

During an outage, don’t risk safety by using unverified splitters, adapters, or cheater plugs—discover what’s truly safe to prevent hazards.

Why LiFePO4 Batteries Matter for Home Backup Planning

A comprehensive look at why LiFePO4 batteries are essential for reliable, safe, and long-lasting home backup systems that safeguard your home.

Pass-Through Charging Explained: Can You Power Devices While Charging?

I’m about to explain pass-through charging and whether you can safely power devices while charging, so read on to find out.