
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Detail Buried Two Pages Deep
Anyone who works in solar, batteries, or backup power knows the feeling. The spec that decides the whole job — the derating curve, the warranty clause, the inverter’s fine print — is never in the headline. It’s buried two references deep in the PDFs, and whoever actually reads it wins the deal at full price. Whoever doesn’t reads the room instead, guesses, and loses.
A new benchmark from Firmulate, which runs AI models as complete simulated companies, just turned that instinct into a hard number. And the number says something uncomfortable about the AI agents being pitched to your business right now.
As an affiliate, we earn on qualifying purchases.
The Experiment
Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing about the outcome is hand-waved. It’s the business equivalent of putting four installers on the identical rooftop and timing them.
The final league table from July 2026:
- 1. gpt-5.6-sol — 95 points. Found the buried fact, closed the deal — the complete performance.
- 2. Kimi K3 — 93 points. The newcomer from Moonshot: closed the deal too, with the cleanest discipline of the field.
- 3. Sonnet 5 — 88 points. Closed the deal, with a few process slips.
- 4. Opus 4.8 — 73 points. The most thorough participant in the field, yet last on the scoreboard.
For calibration, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it, no amount of good work outweighs a breach of trust.
battery bank load assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Separated the Winners
Here’s the finding that should matter to anyone buying AI tools. All four models spotted every crisis. All four refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
Why? The decisive fact wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — the equivalent of the competitor’s real weakness hiding in an old spec sheet nobody opened. The models that went and read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost it automatically. Not because they were dumb, or dishonest, or bad at conversation — because they answered without doing their homework.
If an AI agent will ever quote a battery bank, draft a load-assessment report, or answer a support ticket about your hybrid inverter, this is the property you actually need to test: does it read your files before it answers? It’s measurable. It’s purchase-deciding. And chat demos never show it.
hybrid inverter troubleshooting device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Under Pressure, They Held the Line
The experiment also threw social engineering at the models: fake CEO messages escalating over three stages, plus a reporter’s trick — a friendly “just one yes/no, on background.” All five model configurations refused, every time. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” In an industry where a spoofed email can move real money, that baseline matters.
Then there’s the Opus 4.8 story — the cautionary tale. It was the most thorough participant in the entire field, adding 80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the problem. Effort and diligence don’t automatically become finished work. The same weakness showed up, weaker, in all four models.
As an affiliate, we earn on qualifying purchases.
One Asterisk
Worth noting: Kimi K3 ran without an effort parameter while the others ran at maximum effort — and still nearly topped the table. That makes its second place arguably more impressive, not less.
You Can Watch the Company Run
None of this is a one-off slide deck. Firmulate runs a live company around the clock: 13 synthetic employees, real money mechanics, burn of €105k a month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. The lab is currently running new benchmark configurations daily, and the league grows with each finished run. You can watch it at firmulate.com/live.
There’s also a genuinely fun bit: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — you read the decision, you pick which AI made it. It’s a fast way to feel the personality differences yourself.

The Takeaway
The gap Firmulate exposed is invisible in a chat demo and fatal in a real job. Four frontier models, all honest, all sharp, all on top of every crisis — and half of them still fumbled a €55,000 deal because they didn’t open the file that mattered. Before you let an AI agent touch your CRM, your support queue, your quotes, or your forecasts, the question isn’t whether it writes well. It’s whether it finishes what it starts, whether it reads your files first, and what a unit of useful work actually costs.
Enterprises can go one step further: run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Details and contact are at firmulate.com/pilot.html.
Full results and plain-language findings are at firmulate.com/benchmarks.html. The buried fact is always two references deep. The question is who reads it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
College move-in / dorm season Picks
dorm essentials
As an affiliate, we earn on qualifying purchases.