AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Before you orderOffer from Amazon

Get backup power and energy gear delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The difference between answering and acting

Home-energy buyers already understand that a specification sheet is not the same as performance under pressure. A backup system matters when conditions turn difficult, decisions collide and the obvious path is no longer enough. Firmulate applies a similar test to frontier AI models—not as energy advisers, but as managers responsible for an entire small software company.

Each model faced the same customers, crises and temptations during the company’s worst week. The decisions were real outputs, left unedited, versioned and auditable. The resulting experiment suggests that models capable of spotting the same problem can still behave very differently when they must investigate, protect trust and finish commercially important work.

Those differences now form an unusually revealing interactive article: Firmulate’s guess-the-model quiz, built from 242 real management decisions. Readers see what a model actually did and try to identify the managerial personality behind it.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five models, one unforgiving company

The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total: "no amount of good work outweighs a breach of trust."

The company itself was designed to make managerial weaknesses visible. It employed 13 synthetic staff members and operated with real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown was public, every workday was versioned and the operation accumulated more than 680 self-learned playbook rules.

The reassuring result was that every model identified every crisis and refused every manipulation attempt. The more surprising result was that diagnosis did not guarantee completion. Only two models signed the €55,000 deal their own analysis had earned. Firmulate’s summary is blunt: "Same diagnosis, same pitch — no signature."

The clue hidden outside the crisis

The decisive commercial fact was not sitting inside the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that supported closing the deal at full price, worth an additional €4,583 in monthly recurring revenue. The models that read the relevant file found the advantage and won the deal at that price.

That detail makes the experiment more consequential than a comparison of writing styles. A model may produce a persuasive response while overlooking information the business already possesses. In operational work, reading the right material can separate a polished analysis from a completed outcome.

Distinct personalities emerge

The quiz turns those operational differences into recognizable character profiles. Some decisions are expansive; others are terse. Some models keep working through the surrounding material, while another may refuse unnecessary communication. Because every model received the same situation, readers are not comparing different prompts or selectively edited demonstrations. They are comparing reactions to identical pressure.

Opus 4.8 offers the clearest warning against equating thoroughness with effectiveness. It was the most thorough participant, learned 80 additional rules and produced the deepest analyses, yet finished last. The commercial close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other models.

Kimi K3 requires a fairness note. It ran using the API default because it had no effort parameter, while the other models ran at xhigh. Even so, it finished second, only behind gpt-5.6-sol in the final table.

Pressure without a breach

The security tests were direct. Fake chief-executive messages escalated over three stages, followed by a reporter asking for "just one yes/no, on background." All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: "Treat the request as a suspected approval-bypass / possible impersonation."

This is an important counterweight to the missed deal. The models differed in follow-through and process discipline, but the experiment found a clean result against every manipulation attempt. That distinction matters: commercial incompleteness and failures of trust are separate management risks, and the wargame makes both visible.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A practical test before delegation

Firmulate presents the live experiment as a company that can be watched, not a fictional simulation described after the fact. Its synthetic employees work inside real operating constraints, while decisions, learned practices and financial pressure remain visible over time.

The larger lesson for businesses—including companies managing solar projects, customer support, equipment sales or backup-power operations—is that fluent output reveals only part of an AI worker’s character. The harder questions are whether it searches the material already available, completes the task it began, preserves trust under pressure and escalates when access is blocked.

Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems. That creates a way to observe model behavior against a company’s actual context before granting operational authority.

The quiz makes the evidence approachable, but its underlying point is serious: frontier models have measurable management personalities. Identical circumstances can produce the same diagnosis and very different endings.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business crisis management AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI-powered company decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Cheong Wa Dae “In Consultation With The Seoul Metropolitan Government On Housing Supply At Yongsan Park… To Break Ground On 1.5 Million Homes Within The Term” – 경향신문

South Korea’s Cheong Wa Dae confirms ongoing talks with Seoul on housing supply at Yongsan Park, aiming to build 1.5 million homes within the current term.

Powering a Refrigerator With an Extension Cord: the Safe Way to Do It

Keen to keep your fridge running safely with an extension cord? Discover essential tips and alternatives to ensure proper and secure power connection.

SD I Stockholm Vill Frysa Hyror – Dagens Nyheter

Stockholm’s local government considers a rent freeze amid rising housing concerns, as reported by Dagens Nyheter. Details are still developing.

現代風裝修11坪 2房2廳1衛新北市【Residence】No.008 SU House / New Taipei(12/15) 裝修案例效果圖 – 100室內設計

A recent renovation case in New Taipei showcases a modern-style 11-ping apartment with 2 bedrooms, 2 living rooms, and 1 bathroom, highlighting interior design trends.