
Get backup power and energy gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
You Size Up a Battery Bank Before Trusting It — So Why Not an AI?
Nobody in the home-energy world buys a battery on the vendor’s spec sheet alone. You want the cycle-life data, the depth-of-discharge curve, what happens on a cold night when the grid drops. Independent, adversarial testing under real load. So here’s a question worth sitting with: when a business deploys an AI agent to touch its CRM, support queue, or forecast, what’s the equivalent load test? A chat demo is the spec sheet. A project called Firmulate is building the discharge curve — by making frontier AI models run an actual company through its worst week and scoring what happens. The July 2026 result has a wrinkle nobody predicted: a newcomer from Moonshot beat three of four Western frontier models.
As an affiliate, we earn on qualifying purchases.
The Crucible: Same Company, Same Crisis, Different Brains
Firmulate’s setup is elegantly controlled. Four — in this final run, five — frontier AI models were each handed the same small software company and told to steer it through the same disastrous week: same customers, same crises, same temptations to cut corners. Only the model changes. Every decision is versioned and auditable, so you can go back and see exactly who did what.
The final Crucible league table reads: gpt-5.6-sol in first at 95, Moonshot’s Kimi K3 second at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own framing puts it: no amount of good work outweighs a breach of trust.
enterprise AI decision validation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Result Nobody Bet On
K3 — the newcomer — didn’t stumble into second place. It did the complete job: it found the buried security needle hidden two document references deep in the company’s own files, a fact that sat not in the customer event but in the paperwork the model had to actually read. That buried fact was the decisive competitor weakness, and the models that dug it out won the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue.
K3 also saved the churning customer, and it resisted all three social-engineering baits thrown at it: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused every manipulation attempt, but K3’s on-record reasoning stands out: “Treat the request as a suspected approval-bypass / possible impersonation.”
And here’s the discipline metric that separates the field: across the whole week, K3 logged just one deviation — the cleanest discipline of any model in the experiment.
AI security and trust testing solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Gap That Demos Can’t Show
The broader finding is the uncomfortable part. Every model spotted every crisis. Every model refused every manipulation. Yet only two of the five signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a chat demo. It only shows up under load, which is precisely the point.
Then there’s Opus 4.8, the cautionary tale. It was the most thorough participant — 80 additional learned rules, the deepest analyses in the field — and it still finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, more mildly, in all four other models. Effort and thoroughness, it turns out, are not the same as finishing the job.

As an affiliate, we earn on qualifying purchases.
The Home-Energy Lesson
If you’ve ever compared inverter datasheets, you already understand why this matters. Two batteries can carry identical nameplate specs and behave completely differently under a real load. Two AI models can write equally fluent emails and behave completely differently when the week goes wrong. The league is open now — a newcomer at its API-default effort setting took second place, three spots ahead of one of the most thorough Western models available.
That means picking a model without running your own test is no longer a decision; it’s a bet. Firmulate is doing something about that. The company behind the experiment is live and watchable: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, browse the full benchmark results, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
