AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.
STUDENTS

Prime for Young Adults — start your free trial

Fast free delivery, streaming and member deals for eligible 18–24 year olds.

Try it free

As an affiliate, we earn on qualifying purchases.

The difference between answering and acting

Home-energy buyers already understand that a specification sheet is not the same as performance under pressure. A backup system matters when conditions turn difficult, decisions collide and the obvious path is no longer enough. Firmulate applies a similar test to frontier AI models—not as energy advisers, but as managers responsible for an entire small software company.

Each model faced the same customers, crises and temptations during the company’s worst week. The decisions were real outputs, left unedited, versioned and auditable. The resulting experiment suggests that models capable of spotting the same problem can still behave very differently when they must investigate, protect trust and finish commercially important work.

Those differences now form an unusually revealing interactive article: Firmulate’s guess-the-model quiz, built from 242 real management decisions. Readers see what a model actually did and try to identify the managerial personality behind it.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five models, one unforgiving company

The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total: "no amount of good work outweighs a breach of trust."

The company itself was designed to make managerial weaknesses visible. It employed 13 synthetic staff members and operated with real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown was public, every workday was versioned and the operation accumulated more than 680 self-learned playbook rules.

The reassuring result was that every model identified every crisis and refused every manipulation attempt. The more surprising result was that diagnosis did not guarantee completion. Only two models signed the €55,000 deal their own analysis had earned. Firmulate’s summary is blunt: "Same diagnosis, same pitch — no signature."

The clue hidden outside the crisis

The decisive commercial fact was not sitting inside the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that supported closing the deal at full price, worth an additional €4,583 in monthly recurring revenue. The models that read the relevant file found the advantage and won the deal at that price.

That detail makes the experiment more consequential than a comparison of writing styles. A model may produce a persuasive response while overlooking information the business already possesses. In operational work, reading the right material can separate a polished analysis from a completed outcome.

Distinct personalities emerge

The quiz turns those operational differences into recognizable character profiles. Some decisions are expansive; others are terse. Some models keep working through the surrounding material, while another may refuse unnecessary communication. Because every model received the same situation, readers are not comparing different prompts or selectively edited demonstrations. They are comparing reactions to identical pressure.

Opus 4.8 offers the clearest warning against equating thoroughness with effectiveness. It was the most thorough participant, learned 80 additional rules and produced the deepest analyses, yet finished last. The commercial close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other models.

Kimi K3 requires a fairness note. It ran using the API default because it had no effort parameter, while the other models ran at xhigh. Even so, it finished second, only behind gpt-5.6-sol in the final table.

Pressure without a breach

The security tests were direct. Fake chief-executive messages escalated over three stages, followed by a reporter asking for "just one yes/no, on background." All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: "Treat the request as a suspected approval-bypass / possible impersonation."

This is an important counterweight to the missed deal. The models differed in follow-through and process discipline, but the experiment found a clean result against every manipulation attempt. That distinction matters: commercial incompleteness and failures of trust are separate management risks, and the wargame makes both visible.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A practical test before delegation

Firmulate presents the live experiment as a company that can be watched, not a fictional simulation described after the fact. Its synthetic employees work inside real operating constraints, while decisions, learned practices and financial pressure remain visible over time.

The larger lesson for businesses—including companies managing solar projects, customer support, equipment sales or backup-power operations—is that fluent output reveals only part of an AI worker’s character. The harder questions are whether it searches the material already available, completes the task it began, preserves trust under pressure and escalates when access is blocked.

Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems. That creates a way to observe model behavior against a company’s actual context before granting operational authority.

The quiz makes the evidence approachable, but its underlying point is serious: frontier models have measurable management personalities. Identical circumstances can produce the same diagnosis and very different endings.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business crisis management AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI-powered company decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Chinese PV Industry Brief: Daqo expands beyond polysilicon

Daqo New Energy invests CNY 6 billion in Kunshan to develop advanced energy storage and power equipment for AI data centers, marking a strategic diversification.

We’ve Used This Housekeeping Rule In My Family For 75 Years — It’s The Secret To Spotless Kitchen Drawers

A family shares a 75-year-old housekeeping rule that keeps their kitchen drawers spotless, offering insights into long-standing household routines.

How Inverter Chargers Fit Into Larger Home Backup Systems

Meta Description: Means of seamlessly integrating power sources, inverter chargers are vital, but understanding their full role in home backup systems reveals even more essential functions.

The Free Tool That Preps You For Exams — And Your First Job Interview

A new free online tool is gaining attention for helping students prepare for exams and first job interviews, amid rising search interest.