Only three AI models finished above starting capital in a 500-day startup survival test
Princeton researchers find that most AI agents fail to manage a simulated startup, often underperforming simple rule-based heuristics.
The CEO-Bench benchmark evaluates AI agents on long-horizon tasks by simulating 500 days of startup operations. Results show that current models struggle with strategic decision-making, with only three models maintaining positive capital. The study suggests that modern AI lacks the high-level strategic steering required for complex business management.