AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Anyone who has shopped for solar panels or a home battery knows the difference between a good brochure and a good installer. The brochure has all the answers. The installer has to show up on the roof at 7 a.m., notice that the inverter firmware shipped with the wrong region setting, explain it honestly, and still close out the job. Home energy is a sector where the paperwork is easy and the week-of-installation chaos is hard — and where the difference between the two is money.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

Now ask the same question about AI. The industry currently answers it with chat demos and coding leaderboards: models that write elegant code and answer questions fluently. But if an AI agent will ever touch your installer’s scheduling queue, your support tickets, or your quote pipeline, the real question isn’t whether it writes well. It’s whether it finishes what it starts, reads your files first, and stays honest under pressure.

That is the gap a live experiment at Firmulate set out to measure — and the results should make anyone planning to deploy AI agents in a real business sit up.

Same Company, Same Worst Week, Four Different AIs

Firmulate ran four frontier AI models through an identical scenario: each one was put in charge of the same small software company during its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing rested on a cherry-picked transcript.

The final league table from July 2026 tells a story no chat arena would surface:

  • 1. gpt-5.6-sol — 95 points. The complete performance.
  • 2. Kimi K3 — 93. The newcomer (from Moonshot), with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77 and 5. Opus 4.8 — 73.

For calibration: a do-nothing baseline scores 26. Partial progress counts, but a single breach of trust caps the total — in this scoring philosophy, no amount of good work outweighs a breach of trust. One caveat for fairness: Kimi K3 ran at its API default effort setting while the others ran at high effort, which makes its second-place finish even more striking.

Amazon

solar panel installation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding: Diagnosis Without the Signature

Here is the headline result. All four models spotted every crisis. All four refused every manipulation attempt. And yet only two of them signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.

That gap is invisible in chat demos. A model can identify the opportunity, draft the perfect proposal, and still never finish the job. In home energy terms: it’s the difference between an AI that correctly tells you your battery dispatch schedule is misconfigured and one that actually fixes the ticket and closes it out.

Amazon

home battery system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

What separated the winners from the rest was not intelligence but diligence. The decisive competitive weakness — the fact that made the €55,000 deal closeable at full price — sat two document references deep in the company’s own files. It wasn’t in the customer event at all. The models that read their own company’s documents first won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The models that didn’t, didn’t.

The lesson generalizes painfully well to solar and home energy businesses, which run on dense documentation: interconnection rules, rebate programs, panel spec sheets, permitting paperwork. An AI agent that won’t read the files is a liability regardless of how well it converses.

Amazon

solar inverter firmware update kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Social Engineering Test

The week also included staged manipulations: fake CEO messages that escalated over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five participating models refused. Kimi K3’s on-record reasoning was heartening: “Treat the request as a suspected approval-bypass / possible impersonation.” Under pressure, the field stayed honest.

Amazon

solar project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hardest Lesson: Thorough Isn’t the Same as Good

The most striking profile belongs to Opus 4.8. It was the most thorough participant — over 80 learned rules added during the run, the deepest analyses in the field. And it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, more mildly, in all four models. Effort and care don’t automatically convert into completed work.

You Can Watch the Company Lose Money

Firmulate isn’t a slide deck. The underlying company is real software running every business day: 13 synthetic employees, real money mechanics — burning €105,000 a month against €2,300 in MRR — with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it live at firmulate.com. There’s also a “guess the model” quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The home energy industry is heading toward AI agents that schedule installs, answer support tickets, and draft quotes. Before you trust one, ask the Firmulate question: not “how well does it chat?” but “does it read the files, refuse the scam, and actually sign the deal?” The July 2026 results show those are three different skills — and the market currently only measures the first. Full results and plain-language findings are at firmulate.com/benchmarks.html. Management quality, not chat quality, is the number that will matter when the agent touches your business.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The New IKEA Storage Find That Makes Any Bathroom Look Like It Belongs In A Magazine

A new IKEA storage solution is trending for transforming bathrooms into magazine-worthy spaces, sparking widespread interest and speculation.

Why Smart Battery Monitors Make Backup Planning Easier

Absolutely, smart battery monitors simplify backup planning by providing real-time insights and predictive analytics that help prevent failures before they happen.

A Public Cash Countdown Is Putting AI Management to the Test

A live AI-run software company exposes its cash countdown, workday decisions and management failures—a revealing test of operational trust under pressure.

LFP Vs NMC Batteries: the Tradeoffs No One Explains Clearly

Beneath the surface of LFP and NMC batteries lies a complex tradeoff that could dramatically impact your decision—discover the details that no one clearly explains.