AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The difference between answering and acting

Home-energy buyers already understand that a specification sheet is not the same as performance under pressure. A backup system matters when conditions turn difficult, decisions collide and the obvious path is no longer enough. Firmulate applies a similar test to frontier AI models—not as energy advisers, but as managers responsible for an entire small software company.

Each model faced the same customers, crises and temptations during the company’s worst week. The decisions were real outputs, left unedited, versioned and auditable. The resulting experiment suggests that models capable of spotting the same problem can still behave very differently when they must investigate, protect trust and finish commercially important work.

Those differences now form an unusually revealing interactive article: Firmulate’s guess-the-model quiz, built from 242 real management decisions. Readers see what a model actually did and try to identify the managerial personality behind it.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five models, one unforgiving company

The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total: "no amount of good work outweighs a breach of trust."

The company itself was designed to make managerial weaknesses visible. It employed 13 synthetic staff members and operated with real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown was public, every workday was versioned and the operation accumulated more than 680 self-learned playbook rules.

The reassuring result was that every model identified every crisis and refused every manipulation attempt. The more surprising result was that diagnosis did not guarantee completion. Only two models signed the €55,000 deal their own analysis had earned. Firmulate’s summary is blunt: "Same diagnosis, same pitch — no signature."

The clue hidden outside the crisis

The decisive commercial fact was not sitting inside the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that supported closing the deal at full price, worth an additional €4,583 in monthly recurring revenue. The models that read the relevant file found the advantage and won the deal at that price.

That detail makes the experiment more consequential than a comparison of writing styles. A model may produce a persuasive response while overlooking information the business already possesses. In operational work, reading the right material can separate a polished analysis from a completed outcome.

Distinct personalities emerge

The quiz turns those operational differences into recognizable character profiles. Some decisions are expansive; others are terse. Some models keep working through the surrounding material, while another may refuse unnecessary communication. Because every model received the same situation, readers are not comparing different prompts or selectively edited demonstrations. They are comparing reactions to identical pressure.

Opus 4.8 offers the clearest warning against equating thoroughness with effectiveness. It was the most thorough participant, learned 80 additional rules and produced the deepest analyses, yet finished last. The commercial close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other models.

Kimi K3 requires a fairness note. It ran using the API default because it had no effort parameter, while the other models ran at xhigh. Even so, it finished second, only behind gpt-5.6-sol in the final table.

Pressure without a breach

The security tests were direct. Fake chief-executive messages escalated over three stages, followed by a reporter asking for "just one yes/no, on background." All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: "Treat the request as a suspected approval-bypass / possible impersonation."

This is an important counterweight to the missed deal. The models differed in follow-through and process discipline, but the experiment found a clean result against every manipulation attempt. That distinction matters: commercial incompleteness and failures of trust are separate management risks, and the wargame makes both visible.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A practical test before delegation

Firmulate presents the live experiment as a company that can be watched, not a fictional simulation described after the fact. Its synthetic employees work inside real operating constraints, while decisions, learned practices and financial pressure remain visible over time.

The larger lesson for businesses—including companies managing solar projects, customer support, equipment sales or backup-power operations—is that fluent output reveals only part of an AI worker’s character. The harder questions are whether it searches the material already available, completes the task it began, preserves trust under pressure and escalates when access is blocked.

Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems. That creates a way to observe model behavior against a company’s actual context before granting operational authority.

The quiz makes the evidence approachable, but its underlying point is serious: frontier models have measurable management personalities. Identical circumstances can produce the same diagnosis and very different endings.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business crisis management AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI-powered company decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Battery Cycle Life Explained: Translate Marketing Into Real Years

Meta description: “Many believe battery cycle life numbers tell the full story—discover how real-world habits can extend your battery’s lifespan beyond the specs.

How to Build a Backup Battery Corner That Stays Clean and Safe

Just follow these essential steps to create a safe, clean backup battery corner that ensures long-term reliability and safety.

Running a Refrigerator on Battery Backup: The Setup That Works Overnight

Unlock the secrets to running your refrigerator on battery backup overnight with this proven setup that ensures reliability and safety.

Pass-Through Charging Explained: Can You Power Devices While Charging?

I’m about to explain pass-through charging and whether you can safely power devices while charging, so read on to find out.