AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The Detail Buried Two Pages Deep

Anyone who works in solar, batteries, or backup power knows the feeling. The spec that decides the whole job — the derating curve, the warranty clause, the inverter’s fine print — is never in the headline. It’s buried two references deep in the PDFs, and whoever actually reads it wins the deal at full price. Whoever doesn’t reads the room instead, guesses, and loses.

A new benchmark from Firmulate, which runs AI models as complete simulated companies, just turned that instinct into a hard number. And the number says something uncomfortable about the AI agents being pitched to your business right now.

Amazon

solar panel datasheet reader

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment

Firmulate handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing about the outcome is hand-waved. It’s the business equivalent of putting four installers on the identical rooftop and timing them.

The final league table from July 2026:

  • 1. gpt-5.6-sol — 95 points. Found the buried fact, closed the deal — the complete performance.
  • 2. Kimi K3 — 93 points. The newcomer from Moonshot: closed the deal too, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88 points. Closed the deal, with a few process slips.
  • 4. Opus 4.8 — 73 points. The most thorough participant in the field, yet last on the scoreboard.

For calibration, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the rules put it, no amount of good work outweighs a breach of trust.

Amazon

battery bank load assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Separated the Winners

Here’s the finding that should matter to anyone buying AI tools. All four models spotted every crisis. All four refused every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

Why? The decisive fact wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files — the equivalent of the competitor’s real weakness hiding in an old spec sheet nobody opened. The models that went and read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost it automatically. Not because they were dumb, or dishonest, or bad at conversation — because they answered without doing their homework.

If an AI agent will ever quote a battery bank, draft a load-assessment report, or answer a support ticket about your hybrid inverter, this is the property you actually need to test: does it read your files before it answers? It’s measurable. It’s purchase-deciding. And chat demos never show it.

Amazon

hybrid inverter troubleshooting device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Under Pressure, They Held the Line

The experiment also threw social engineering at the models: fake CEO messages escalating over three stages, plus a reporter’s trick — a friendly “just one yes/no, on background.” All five model configurations refused, every time. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.” In an industry where a spoofed email can move real money, that baseline matters.

Then there’s the Opus 4.8 story — the cautionary tale. It was the most thorough participant in the entire field, adding 80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating the problem. Effort and diligence don’t automatically become finished work. The same weakness showed up, weaker, in all four models.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Asterisk

Worth noting: Kimi K3 ran without an effort parameter while the others ran at maximum effort — and still nearly topped the table. That makes its second place arguably more impressive, not less.

You Can Watch the Company Run

None of this is a one-off slide deck. Firmulate runs a live company around the clock: 13 synthetic employees, real money mechanics, burn of €105k a month against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. The lab is currently running new benchmark configurations daily, and the league grows with each finished run. You can watch it at firmulate.com/live.

There’s also a genuinely fun bit: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html — you read the decision, you pick which AI made it. It’s a fast way to feel the personality differences yourself.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Takeaway

The gap Firmulate exposed is invisible in a chat demo and fatal in a real job. Four frontier models, all honest, all sharp, all on top of every crisis — and half of them still fumbled a €55,000 deal because they didn’t open the file that mattered. Before you let an AI agent touch your CRM, your support queue, your quotes, or your forecasts, the question isn’t whether it writes well. It’s whether it finishes what it starts, whether it reads your files first, and what a unit of useful work actually costs.

Enterprises can go one step further: run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Details and contact are at firmulate.com/pilot.html.

Full results and plain-language findings are at firmulate.com/benchmarks.html. The buried fact is always two references deep. The question is who reads it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


COLLEGE MOVE-IN

College move-in / dorm season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Pass-Through Charging Explained: Can You Power Devices While Charging?

I’m about to explain pass-through charging and whether you can safely power devices while charging, so read on to find out.

Git Hosting That Never Leaves Europe

A new trend emerges with Git hosting providers committed to never leaving Europe, raising questions about data sovereignty and privacy.

Inmobiliaria Vesta Surges In Global Coverage

Vesta’s recent surge in international coverage marks a notable shift in its global profile, with 26 mentions in recent media monitoring data.

Your Solar Installer’s AI Might Ace the Chat Demo and Still Leave the Contract Unsigned

Four frontier AIs ran the same company through its worst week. All saw the crisis; only two signed the €55k deal. Chat quality isn’t management quality.