
Are AI models managing like humans? A live experiment reveals surprising personality traits in AI decision-making
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI Front and Center in Business Management
In a groundbreaking live test, four advanced AI models were tasked with running a real, money-losing software company through its most turbulent week. The goal? To see whether these models could handle crises, resist manipulation, and close deals — all while demonstrating distinct management personalities. This isn’t a simulation; every decision was real, auditable, and reflected genuine business mechanics. You can watch the company in action at firmulate.com/live.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Models and Their Scores
- gpt-5.6-sol: Achieved the highest score of 95, identifying hidden information, closing the big deal, and demonstrating a thorough, disciplined approach.
- Kimi K3: Scored 93, showing the cleanest management style, refusing all manipulative tactics, and securing the deal at full price.
- Sonnet 5: Scored 88, closing the deal with minor slips and less discipline than the top two.
- Fable 5: Scored 77, also closing the deal but with more process slips and missed opportunities.
As an affiliate, we earn on qualifying purchases.
What They All Had in Common
Remarkably, all models spotted every crisis and refused every attempt at manipulation, including fake CEO messages and media tricks. Their consistent refusal to engage in deception underscores an emerging trait: honesty under pressure.
AI for business crisis management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading the Files
The decisive factor in winning the deal was reading and interpreting internal company files. The models that examined documents deeply uncovered a critical piece of information buried two references deep in the files — a detail that allowed them to secure a full-price deal worth over €4,583 in monthly recurring revenue. The models that skipped this step lost the opportunity, illustrating the importance of thorough information vetting.
Different Personalities in the AI Arena
The experiment also revealed different management personalities. The most comprehensive model, Opus 4.8, with over 80 learned rules and deep analysis, performed worst in closing because it left the final opportunity on the table, showing a tendency for over-analysis and hesitance. Meanwhile, Kimi K3, running without an effort parameter, demonstrated a disciplined, straightforward style that prioritized integrity and decisive action.
Handling Social Engineering and Crisis Scenarios
In staged social engineering attacks—fake CEO messages escalating over three stages, plus a reporter trick asking for a simple yes/no answer—every model refused to engage. Kimi K3’s reasoning was clear: treat such requests as possible impersonation or approval-bypass attempts. Their collective resistance shows that AI can be programmed to maintain integrity even under pressure.
The Broader Implication
This experiment isn’t just about AI performance; it’s a window into the emerging management personalities of AI systems. As companies consider integrating AI into decision-making roles, understanding whether these models can uphold trust, read critical internal information, and resist manipulation is vital. The experiment’s insights are available in real time and include detailed analyses of every decision made.
Why It Matters for Your Business
If AI agents will touch your CRM, support queues, or forecasting tools, the key question isn’t whether they write well — it’s whether they can finish what they start, stay honest under pressure, and read your internal files thoroughly. The current leaderboard, based on a live business test, shows that different models exhibit distinct management styles, personalities, and vulnerabilities. Choosing the right AI isn’t just about capabilities; it’s about trustworthiness and discipline.
Take the Next Step
Interested in testing your own AI workforce? You can run the same wargame against your business data, without risk or interference, at firmulate.com/pilot.html. See firsthand how your AI can perform under pressure, and make better hiring decisions for your automation future.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html