AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

You Size Up a Battery Bank Before Trusting It — So Why Not an AI?

Nobody in the home-energy world buys a battery on the vendor’s spec sheet alone. You want the cycle-life data, the depth-of-discharge curve, what happens on a cold night when the grid drops. Independent, adversarial testing under real load. So here’s a question worth sitting with: when a business deploys an AI agent to touch its CRM, support queue, or forecast, what’s the equivalent load test? A chat demo is the spec sheet. A project called Firmulate is building the discharge curve — by making frontier AI models run an actual company through its worst week and scoring what happens. The July 2026 result has a wrinkle nobody predicted: a newcomer from Moonshot beat three of four Western frontier models.

Amazon

AI load testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible: Same Company, Same Crisis, Different Brains

Firmulate’s setup is elegantly controlled. Four — in this final run, five — frontier AI models were each handed the same small software company and told to steer it through the same disastrous week: same customers, same crises, same temptations to cut corners. Only the model changes. Every decision is versioned and auditable, so you can go back and see exactly who did what.

The final Crucible league table reads: gpt-5.6-sol in first at 95, Moonshot’s Kimi K3 second at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own framing puts it: no amount of good work outweighs a breach of trust.

Amazon

enterprise AI decision validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Result Nobody Bet On

K3 — the newcomer — didn’t stumble into second place. It did the complete job: it found the buried security needle hidden two document references deep in the company’s own files, a fact that sat not in the customer event but in the paperwork the model had to actually read. That buried fact was the decisive competitor weakness, and the models that dug it out won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue.

K3 also saved the churning customer, and it resisted all three social-engineering baits thrown at it: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused every manipulation attempt, but K3’s on-record reasoning stands out: “Treat the request as a suspected approval-bypass / possible impersonation.”

And here’s the discipline metric that separates the field: across the whole week, K3 logged just one deviation — the cleanest discipline of any model in the experiment.

Amazon

AI security and trust testing solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Gap That Demos Can’t Show

The broader finding is the uncomfortable part. Every model spotted every crisis. Every model refused every manipulation. Yet only two of the five signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a chat demo. It only shows up under load, which is precisely the point.

Then there’s Opus 4.8, the cautionary tale. It was the most thorough participant — 80 additional learned rules, the deepest analyses in the field — and it still finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four other models. Effort and thoroughness, it turns out, are not the same as finishing the job.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI model evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Home-Energy Lesson

If you’ve ever compared inverter datasheets, you already understand why this matters. Two batteries can carry identical nameplate specs and behave completely differently under a real load. Two AI models can write equally fluent emails and behave completely differently when the week goes wrong. The league is open now — a newcomer at its API-default effort setting took second place, three spots ahead of one of the most thorough Western models available.

That means picking a model without running your own test is no longer a decision; it’s a bet. Firmulate is doing something about that. The company behind the experiment is live and watchable: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, browse the full benchmark results, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mortgage Rates Today

Mortgage rates in the US have increased modestly today, reflecting ongoing market volatility. Details remain developing, with broader trends still uncertain.

Why Power Distribution Boxes Help Temporary Backup Feel Safer

Guiding you through safety features and proper setup, power distribution boxes help make temporary backup power feel safer and more reliable during outages.

How to Use a Power Station Safely Indoors (Ventilation Myths Included)

No matter the myth, proper ventilation is crucial when using a power station indoors—discover essential safety tips to keep you protected.

10 Home Deals From T.J. Maxx You Don’t Want To Miss This Labor Day

Discover the top 10 home deals from T.J. Maxx this Labor Day, offering significant savings on furniture, decor, and essentials. Limited-time offers you can’t miss.