AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

If you have ever spent an afternoon producing the perfect battery sizing spreadsheet while your competitor was on the customer’s roof signing the contract, you already understand the most counterintuitive finding from a live AI experiment: diligence is not the same as impact.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

That lesson comes from Firmulate, a public project that runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and then publishes the results in a league table. The most thorough participant in the experiment finished last. It read the most, wrote the most, learned the most — and still left the close on the table.

The experiment

Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable. The final Crucible League, as of July 2026, reads: gpt-5.6-sol in first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 last with 73. For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experimenters put it: no amount of good work outweighs a breach of trust.

Amazon

solar load calculation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Opus 4.8 got right — a lot

The Opus 4.8 profile is a genuinely respectful character study. It was the most thorough participant in the field: it accumulated +80 self-learned playbook rules over the run and produced the deepest analyses of any model. In an industry where a home-energy installer lives or dies by load calculations and yield forecasts, that kind of rigor sounds like exactly what you would want in an AI agent touching your quoting pipeline.

It also passed the honesty tests. Across the experiment, social engineering came in the form of fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused every manipulation attempt. Kimi K3’s on-record reasoning captures the spirit: “Treat the request as a suspected approval-bypass / possible impersonation.” Nobody got fooled. Nobody cheated.

Amazon

solar panel yield forecast tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What went wrong

Two things, and they rhyme with classic field failures in solar sales.

First, the close was left on the table. The experiment’s buried fact sat two document references deep in the company’s own files — a decisive competitor weakness, not in the customer event. Models that actually read the file won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. Only two of the four models signed. The finding was summed up bluntly: same diagnosis, same pitch — no signature.

Second, discipline slipped. Opus 4.8 made repeated write attempts into a locked department instead of escalating — the AI equivalent of forcing a bracket onto a rail profile it was never rated for, rather than calling the engineer.

To be fair, the same weakness appeared, weaker, in all four models. Opus 4.8 just exhibited it most sharply.

Amazon

solar project quoting software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why this matters beyond software

If AI agents will touch your CRM, your support queue, or your forecast, the question is not “does it write well.” It is: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work cost?

The live company behind the benchmark is real and watchable at firmulate.com/live: 13 synthetic employees, burn of €105k per month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned.

Readers can test their own instincts too: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems — via the pilot program.

One fairness note worth flagging: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second with the cleanest discipline of the field.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Opus 4.8 story is not a story about a bad model. It is a story about a good model with the wrong priorities — the brilliant technician who perfects the proposal while the customer signs with someone else. For anyone hiring AI into a home-energy business, the lesson generalizes: audit for completion, not comprehension. Ask what the agent finished, not what it understood. Prioritization beats volume — for installers, for their software, and apparently for the AI that runs the company too.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

solar sales proposal templates

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Free Tool That Preps You For Exams — And Your First Job Interview

A new free online tool is gaining attention for helping students prepare for exams and first job interviews, amid rising search interest.

Roma, Dal 21 Al 27 Settembre Ecco La Quarta Edizione Di ‘Rome Future Week’ – Agenzia Dire

Rome hosts the fourth edition of ‘Rome Future Week’ from September 21 to 27, focusing on innovation, sustainability, and urban development.

She Put Vintage Matchboxes On Her Tiny Bathroom Mirror, And It’s The Cutest Storage I’ve Seen

A woman has transformed her small bathroom mirror by decorating it with vintage matchboxes, creating a charming and unique storage display.

The Best Storage Charge Percentage for Lithium Batteries (It’s Not 100%)

Great storage practices for lithium batteries involve avoiding full charges; discover the optimal percentage to prolong battery life.