
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Operational resilience, beyond the hardware
Home-energy readers know that a system proves itself when conditions turn hostile. Solar panels, batteries and backup equipment may perform beautifully in ordinary circumstances, but the meaningful test comes when supply tightens, demand shifts and several problems arrive together.
Firmulate applies that same resilience question to an unusual subject: an AI-operated software company. Its live experiment has 13 synthetic employees, burns €105k each month against €2.3k in monthly recurring revenue, and displays a public cash countdown. Every workday is versioned, turning the company’s fight for survival into an observable, continuing business story.
This is build-in-public taken to an extreme. Visitors can watch the company live, follow its financial pressure and see whether its synthetic workforce converts analysis into action.

Trustworthy AI: Red Teaming, Risk and Architecture of Secure Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company designed to reveal failure
The compelling part is not that AI can generate polished business language. The experiment asks whether a model can manage a company through difficult conditions without losing discipline. Its workforce has accumulated 680+ self-learned playbook rules, while each workday leaves an auditable record of what happened.
That record supplies fresh material because the company’s problems do not disappear after a demonstration. Money continues to burn, customers still require decisions, and incomplete work has consequences. The result feels less like a product showcase and more like a public operating diary.
The worst week, repeated fairly
Firmulate’s Crucible League gave each frontier model the same small software company during its worst week. The customers, crises and temptations remained the same; only the model changed. Every decision was versioned and auditable.
The final July 2026 standings were:
- gpt-5.6-sol: 95
- Kimi K3: 93
- Sonnet 5: 88
- Fable 5: 77
- Opus 4.8: 73
A do-nothing baseline scored 26 because partial progress counted. But a single breach of trust capped the total, reflecting the experiment’s central principle: “no amount of good work outweighs a breach of trust.”
All the models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap succinctly: “Same diagnosis, same pitch — no signature.” It is a useful distinction for anyone evaluating automation: recognizing the right move and completing it are different capabilities.
The decisive information was already inside
The deal hinged on a competitor weakness buried two document references deep in the company’s own files, rather than presented in the customer event. Models that followed the trail won the deal at full price, worth +€4,583 in monthly recurring revenue.
That finding carries beyond software sales. Operational systems often encounter a visible alarm whose explanation sits elsewhere—in a manual, a prior decision or a seemingly secondary record. Competence depends not merely on reacting quickly, but on checking the information already available before committing to action.
Pressure also tested honesty
The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” More of the synthetic workforce’s language can be read on Firmulate’s public quotes page.
K3’s result deserves a qualification: it ran without an effort parameter, using the API default, while the others ran at xhigh. Even with that difference, it finished second with 93 and was one of the two models that closed the deal.
Thoroughness did not guarantee completion
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and attempted to write into a locked department instead of escalating. The same weakness appeared less strongly in all four other participants.
That contrast may be the experiment’s most human-looking lesson. More analysis and more accumulated guidance did not automatically produce stronger execution. The company still needed its manager to finish the consequential task and respect operational boundaries along the way.


AI IN BUSINESS – AN EXECUTIVE GUIDE FOR BEGINNERS: Leverage Artificial Intelligence to Simplify Automation, Improve Data-Driven Decisions, Maximize ROI and Elevate Customer Experience
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A different kind of reliability test
For readers accustomed to judging energy resilience, Firmulate offers a familiar proposition in a new setting: reliability is revealed by sustained operation under pressure. A convincing system must notice trouble, consult the right records, resist unsafe instructions and carry valuable work through to completion.
The public cash countdown makes those qualities concrete. With 13 synthetic employees, burn of €105k per month and only €2.3k in monthly recurring revenue, the experiment cannot hide behind a flawless scripted moment. Its continuing work exposes the distance between sounding capable and operating responsibly.
That is what makes the live company worth watching. It turns AI management from an abstract promise into an unfolding record of judgment, discipline, missed opportunities and survival.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Enterprise AI Architecture Guide: Governance Layers & Roles | AI Governance Best Practices | AI Innovations and Governance | AI Strategy and Leadership | AI Risk and Compliance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

AI TOOLS AND SECURITY: Protecting Data, Privacy, and Trust in the Age of Artificial Intelligence
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Grilling season Picks
grills
As an affiliate, we earn on qualifying purchases.