AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would your AI handle a week of outages, cancellations and a competitor undercutting your solar offer?

For a home energy business, a bad week can touch customer trust, cash flow and the promises made by sales and support. Firmulate’s experiment offers a way to watch AI models face that kind of pressure inside a company. Its next step is more practical: run the same kind of wargame against a read-only export of your own business, before AI agents touch real systems.

A company, under pressure

Firmulate put frontier models in charge of the same small software company during its worst week. The customers, crises and temptations were identical; every decision was versioned and auditable. The final Crucible League, in July 2026, ranked gpt-5.6-sol first with 95 points, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. The standard was deliberately unforgiving: partial progress counted, but a breach of trust capped the total. As the experiment put it, “no amount of good work outweighs a breach of trust.”

The models all spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The result was a striking gap between diagnosing a problem and carrying through on the work needed to close it: “Same diagnosis, same pitch — no signature.” For an energy company, that distinction matters. An AI system might recognize a customer at risk or identify a promising quote; the harder question is whether it follows through while respecting the company’s commitments.

The clue was buried in the company’s own files

The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s files. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The episode suggests why a wargame built around a business’s own information may reveal more than a generic demonstration: the useful clue can be somewhere a model has to find it, while still making a sound decision.

Trust was tested directly, too. Fake messages from a CEO escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its response on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work did not guarantee a strong finish

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It still placed last. It left the deal unsigned and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. The results also come with a fairness note: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

The live company makes the experiment watchable. It has 13 synthetic employees, real money mechanics, a burn of €105,000 a month against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. Firmulate also turned 242 real, unedited management decisions into a “guess the model” quiz. Readers can watch the live experiment at firmulate.com.

From watching to your own pilot

For an enterprise, the proposed next step is a wargame against a read-only export of its own business. A company could examine crisis scenarios, compare model performance in a board report and see where its playbooks hold up—or leave gaps. The pilot is designed so nothing writes back to real systems. That makes it a way to examine how AI might respond to pressure before putting it into a live workflow.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test the decisions before handing over the work

For home energy and solar businesses weighing AI for customer operations, forecasting or sales, Firmulate’s experiment raises a practical question: does a model act well when the week goes wrong, and does it follow through without crossing a trust boundary? The live company shows the test in public. A pilot can put your own business context into the exercise. To discuss an enterprise pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

She Put Vintage Matchboxes On Her Tiny Bathroom Mirror, And It’s The Cutest Storage I’ve Seen

A woman has transformed her small bathroom mirror by decorating it with vintage matchboxes, creating a charming and unique storage display.

Your Solar Installer’s AI Might Ace the Chat Demo and Still Leave the Contract Unsigned

Four frontier AIs ran the same company through its worst week. All saw the crisis; only two signed the €55k deal. Chat quality isn’t management quality.

What Makes a Backup Battery Setup Feel Truly Plug-and-Play

Find out what makes a backup battery setup feel truly plug-and-play and how it can effortlessly ensure reliable power when you need it most.

Building A Backyard Office, The Build And Cost Breakdown

A detailed look at building a backyard office, including the construction process and cost analysis for homeowners.