ConceptioArchivearXiv CS
arXiv CSopen access

CEO-Bench: Can Agents Play the Long Game?

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

CE O - B E N C H : Can Agents Play the Long Game? Haozhe Chen

Karthik Narasimhan

Zhuang Liu

arXiv:2606.18543v1 [cs.AI] 16 Jun 2026

Princeton University

Code Business database Customers Ledger R&D projects

Subscriptions Ad leads More tables

Company management tools MONEY GROWTH PRODUCT SALES

Spend Ads R&D Deals

Prices Promo Social Capacity Market More

Social media

@dana_p @nm_official @founder_li

latency on Pro is unusable… capacity +4× — reliability … finally a billing UX that do…

Blog

▶ Trajectory Market

Language Model Agent

ECONOMIC CYCLE CUSTOMERS Enterprise Individual COMPETITOR

Terminal ANALYZE DATA

d28 +0.09 quality

$ query("SELECT category, SUM(amount) FROM ledger …") $ query("SELECT channel, group, spend/leads AS cpa …") $ query("SELECT segment, churn_rate FROM customers …")

TAKE ACTIONS

Outcomes

$ pricing.set_prices(A=10.0, B=20.0, C=35.5) $ infrastructure.set_capacity_tier(2) $ marketing.post_social_media("capacity +4×") $ research.start_research_project(tier=2)

d66

$ next_week()

d500

ADVANCE SIMULATION

0.42M 2.64M

d70 d71 d82

tickets resolved −312 new subscribers +98 cancellations −184 R&D finished · +0.18 Final cash

$4.21M

Figure 1. CEO-B E N C H evaluates general long-horizon agent capabilities by simulating a startup over 500 days in a realistic and challenging environment. The agent operates through a programmable interface with access to business databases, company management tools, and social media. Outcomes are driven by a partially observable, noisy, and evolving market with delayed and coupled consequences.

Abstract Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a combination of sophisticated skills that remain largely untested in agents: (1) navigating long horizons amid uncertainty; (2) acquiring information in noisy environments; (3) adapting to a changing world; (4) orchestrating multiple moving parts toward a coherent goal. We introduce C E O - B E N C H , which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. An agent manages pricing, marketing, budgeting, and many other aspects of a fictional company through a programmable Python interface, operating in the same environment and facing the same challenges as a human CEO. Success demands analyzing noisy, interconnected business databases, translating signals into sound strategy, and coordinating many decisions with programming. The strongest agents write sophisticated code that simulates customer cohorts to forecast future cash and mines negotiation history to uncover hidden customer preferences. Even so, most state-of-the-art models struggle in this environment. Only Claude Opus 4.8 and GPT-5.5 finish above the $1M starting balance, and neither consistently turns a profit. C E O - B E N C H takes a first step toward measuring the intelligence required to drive sustained, adaptive progress over time. Correspondence to Haozhe Chen at [email protected] and Zhuang Liu at [email protected].

1

$10B

Estimated final cash upper bound ($2.2B)

$1B

Claude Opus 4.8 ($27.8M)

GPT-5.5 ($21.3M)

)gol ,DSU( dnah no hsaC

$100M

Rule-based baseline ($15.76M)

$10M Starting cash ($1M)

$1M

Claude Opus 4.7

$100k

Claude Sonnet 4.6 Kimi K2.6

$10k DeepSeek V4 Pro

$1k $100

0

☠️

Grok 4.20 100

☠️

Gemini 3 Flash Claude Haiku 4.5

200

☠️☠️

Day

300

☠️

GLM 5.1 400

500

Figure 2. Cash on hand over time for each model’s best run. Most state-of-the-art models struggle to complete the simulation without bankruptcy. Only Claude Opus 4.8 and GPT-5.5 grow cash above the $1M starting balance for their best runs. Current models still struggle to combine long-horizon planning, noisy information gathering, adaptation, and coordinated execution over time.

1 Introduction Language model agents are becoming increasingly capable at short-horizon tasks. They can fix a GitHub issue (Jimenez et al., 2024), follow a service policy in dialogue (Yao et al., 2025), or complete a web workflow (Zhou et al., 2024). These are real skills, but they share a simple shape: the agent gets a clear goal, acts for a short time, and receives feedback quickly. As agents approach reliable execution of such individual tasks, a natural next question is what we should expect them to do after the local task is no longer the bottleneck. Human intelligence goes beyond local execution (Newell and Simon, 1972; Simon, 1955). Many consequential human achievements are not single well-specified tasks, but long chains of decisions made under uncertainty: choosing what to prioritize, allocating limited resources, interpreting noisy signals, and adapting as conditions change (Simon, 1955; March, 1991; Teece et al., 1997). Future agents will need the same kind of sustained strategic control if they are to move beyond task completion and operate effectively in the real world. Early agent evaluations such as SWE-bench (Jimenez et al., 2024), WebArena (Zhou et al., 2024), and τbench (Yao et al., 2025) evaluate real-world skills, but they are scoped to short episodes with quickly observed outcomes. GDPval (Patwardhan et al., 2025) broadens evaluation to economically valuable work, but remains a one-shot deliverable rather than a persistent process. Agentic-memory benchmarks test agents’ ability to use information over time, but they primarily measure storage and retrieval skills (Hu et al., 2026; He et al., 2026b). Vending-Bench (Backlund and Petersson, 2025a;b) and Accounting-Bench (Penrose AI, 2025) take a first step toward evaluating agents in long-horizon simulated environments. Yet these settings involve a narrow set of decisions and largely stable environments. They do not test whether agents can coordinate many interdependent actions, acquire information from noisy feedback, and devise strategy amid delayed consequences and changing conditions. The next stage of agent evaluation calls for a shift toward settings where actions accumulate over long horizons; environment state is only observable through indirect evidence; feedback is noisy and delayed; and conditions continue to change. Success depends on more than individual capabilities. Agents must integrate them into coherent behavior, make long-horizon plans, turn accumulated evidence into actionable signals, coordinate interdependent decisions, and continuously adapt strategy as new information arrives. C E O - B E N C H instantiates this challenge in a realistic, large-scale startup simulation, where an agent runs a company for 500 days through a programmable Python interface with 34 tools and a 19-table business database. We show the structure of C E O - B E N C H in Fig. 1. Beyond issuing individual tool calls, the agent writes and executes code, querying the database with SQL to analyze the company’s state and composing the available tools into custom workflows. It thus operates in the same environment and faces the same challenges 2

$100M $10M $1M $100k $10k $1k $100

$100M $10M $1M $100k $10k $1k $100

$100M $10M $1M $100k $10k $1k $100

Claude Opus 4.8

0

100

200

Day

300

400

500

Kimi K2.6

0

100

200

Day

☠️

300

400

500

Claude Haiku 4.5

0

☠️ ☠️ ☠️

100

200

Day

300

400

500

$100M $10M $1M $100k $10k $1k $100

$100M $10M $1M $100k $10k $1k $100

$100M $10M $1M $100k $10k $1k $100

GPT-5.5

0

☠️

100

200

Day

300

400

☠️

500

Claude Sonnet 4.6

0

100

☠️

200

Day

300

☠️

400

500

Gemini 3 Flash

0

100

☠️☠️ ☠️ 200

Day

300

400

500

$100M $10M $1M $100k $10k $1k $100

$100M $10M $1M $100k $10k $1k $100

$100M $10M $1M $100k $10k $1k $100

Claude Opus 4.7

0

100

200

Day

300

400

500

300

400

500

300

400

500

GLM 5.1

☠️

0

100

☠️

200

Day

☠️

DeepSeek V4 Pro

☠️☠️ ☠️

0

100

200

Day

Figure 3. Cash on hand over time for each of the three runs per model. We show all experiments for each model to describe full behavior patterns. We also release all agent action trajectories in an interactive trajectory viewer.

as a human running the company, and the task demands coding and data-analysis skill together with strategic thinking. As shown in Fig. 4, the agent must coordinate diverse operating decisions across pricing, growth, product, operations, communication, enterprise sales, and more. Decisions play out over realistic business timelines: revenue arrives on billing cycles, R&D takes days to weeks, and mistakes surface later through churn and reputation, forcing long-horizon reasoning under uncertainty. Much of the state is hidden, so the agent must infer customer satisfaction, willingness to pay, and shifting preferences from noisy signals in data analytics and social media. At the same time, we design the environment to keep changing as a result of customer preference drift, macroeconomic cycles, and competitor shocks. Success requires continually revising strategy in response to these shifting conditions while maintaining coherent decisions across the business. Our evaluation shows that this challenge remains difficult for agents built on current state-of-the-art models. We show in Fig. 2 that while most agents can produce valid tool calls and analytics queries, they struggle to sustain coherent strategy over time and often bankrupt before completing the simulation. Only Claude Opus 4.8 and GPT-5.5 finish above the $1M starting balance and the non-LLM rule-based baseline. We show in Fig. 3 that even these two models fail to consistently profit across experiments. Our analysis shows that performance correlates with core capabilities targeted by C E O - B E N C H : inferring hidden structure from noisy data, forecasting delayed consequences, and adapting to competitive pressure. By inspecting agent action trajectories, we find distinctive behavior patterns across models. For example, GPT-5.5 and Claude Opus 4.8 actively explores various strategies, while Claude Opus 4.7

Marketing

Growth

Research

Product Quality

Customer Support

Agent

Infrastructure

Analytics

Competitors

Market Shocks

Pricing

Economy

Development

Reputation

Market Research Social Media

Enterprise Sales Cash Flow

Budgeting

Figure 4. Running a startup requires coordinating many moving parts, making it a fitting choice as a canonical task evaluating agent’s skills to steer complex decisions across long-horizon.

3

Category

Actions

Example tools

Database query

Query 19 business SQL databases and conduct data analytics

query

Monetization

Set prices, usage quotas, discounts, and in-product ads

pricing.set_prices, pricing.set_usage_quotas

Growth and market expansion

Allocate targeted advertising spend and promotion across channels and customer groups

marketing.set_targeted_ad_spend, marketing.set_lead_promotion

Product quality and R&D

Choose model tiers, fund day-to-day development, and pricing.set_model_tiers, launch research projects research.start_research_project

Reliability

Buy infrastructure capacity and fund customer support

infrastructure.set_capacity_tier, analytics.set_targeted_ops_spend

Enterprise sales

Conduct multi-turn negotiations over price and plan with enterprise prospects and renewals

enterprise.send_enterprise_deal, enterprise.reject_enterprise_deal

Information acquisition

Pay for market research to discover new customer groups and learn more about existing groups

market.research_market, market.research_group

Public communication

Monitor social media for customer complaints, competitor news, and economic trends, then post or reply to influence growth

marketing.post_social_media, analytics.get_social_posts

Table 1. Agent action space categories and example tools in CEO-B E N C H . These tools enable agents to design diverse operation strategies but also pose challenges on coordinating many actions toward one coherent goal.

limits itself to a passive cash-preservation strategy. Claude Opus 4.8 reaches more customers initially and drop to zero customer mid simulation, while GPT-5.5 maintains consistent customer base throughout. These results show that CEO-B E N C H exposes fine-grained behavioral patterns that remain invisible to existing evaluations, while revealing substantial headroom in models’ ability to integrate individual capabilities into coherent, adaptive behavior over extended horizons.

2 Designing C E O - B E N C H In this section, we provide an overview of how the CEO-B E N C H simulator works. Then, we describe how we design the action interface to make the task open-ended. Finally, we detail world mechanics design considerations that make the task challenging and realistic.

2.1 How CEO-B E N C H Works In CEO-B E N C H , an agent runs a fictional subscription-software company called NovaMind for 500 simulated days. It begins on day one with zero customers and $1M in cash and is graded on cash on hand at the end. If cash ever falls strictly below zero, bankruptcy ends the simulation. We provide an overview of the simulator mechanics in this section and include full details in Appendix A. What an agent can do. For each simulated week, the agent can take actions for unlimited turns across 34 tools in the categories displayed in Table 1. These categories cover pricing and plan design, growth and market expansion, product quality and research, reliability and support, information acquisition, public communication, and enterprise sales. Each tool accepts fine-grained structured arguments, so agents can compose a large space of possible policies. Section 2.3 explains the tool interface design in more detail. How an agent makes and loses money. An agent makes profits through customer subscription payments and in-product ad monetization. We abstract the company product that customers subscribe to as a numerical product quality. Higher product quality results in more product subscriptions and payments, but maintaining quality via development, research, infrastructure capacity, support, and model tier choices requires spending. Acquiring customers through advertising channels also costs money. Cash therefore changes through both immediate costs and delayed revenue effects. We show the calculation of cash change between each day and decompose each contributing factor in Equation 1. We fully explain the role and mechanism of each factor in

4

Appendix A.3, A.4 and A.8. B − Bt = | t+1{z } daily cash change

ads use + ∑ Yi,t − ∑ χ p U p,t − Kκ t − xt − xtdev |{z} |{z} | {z } p i | {z } subscription dev {z } support | capacity usage

capacity

Ytsub |{z}

ops

spending cost spending in-product usage compute cost ads target-ops target-dev group project ads − Xt − x g,t − xc,g,t − Kt − Ntlead clead − Ktmarket − Kt | | {z } {z } | {z } | {z } | {z } g c,g group lead market | {z } | {z } targeted research acquisition research projects research support acquisition targeted ads dev payments

(1)

Modeling customers and indirect feedback. There are 26 customer groups in the simulator. Each customer group consists of a distribution of hidden price and quality preferences, such as a maximum willingness to pay and a minimum accepted quality at each price. Each customer is created by sampling its unique preference parameters from a group distribution. At a subscription plan’s price, a customer subscribes if the offered product quality exceeds the customer’s minimum accepted quality. The customer may switch plans if another plan gives a better quality surplus and may cancel if no plan remains acceptable. We show the mathematical definition of a customer’s price-quality preference curve in Equation 5. We describe full details of each factor in Appendix A, Subsection A.2. Example curve plots appear in Fig. 16. Customer satisfaction changes company reputation, and reputation affects the new customer acquisition rate. The agent does not directly observe satisfaction, willingness to pay, or quality thresholds. It instead infers feedback by analyzing subscription, churn, support, revenue, and reputation data and by monitoring simulated social media. Customer acquisition and enterprise negotiation. Agents acquire new customers by spending on advertising channels. Each customer group reacts differently to each ad channel, so the same spend can produce different acquisition rates across groups. Reputation, social media reactions, market saturation, demand surges, and macroeconomic conditions also affect acquisition speed. We show the calculation of expected new prospective customers for group g on day t in Equation 2. We describe full details of each factor in Appendix A.5 and A.6. We sample daily from a Poisson distribution parameterized by this expectation. Market research can reveal additional customer groups and improve what the agent knows about known groups. Enterprise customers follow the same price and quality logic, but deals are negotiated through offers, counter-offers, reply delays, and possible rejection.  h

prospect

E n g,t | {z

i }

expected new prospective customers for group g

      x L c,g,t c,g,t  net  = R g,t · Dg,t · Ct · Mg,t · A g,t · Zt · ∑ + ∑ Nh,t Wh,g  |{z} |{z}  c  xad |{z} |{z} |{z} |{z} h  {z } | {z }  reputation market saturation calendar macro econ social media demand  |  in group g

for group g

cycle

cycle

reaction

surge

leads from each ad channel

networking effect from each group

(2) Product quality and competitor pressure. Product quality is affected by daily development, research projects, model tier choices, targeted development, infrastructure capacity, support spending, usage quotas, and in-app ad strength. These controls shape customer experience through base product quality, quota saturation, system outages, support delays, relationship history, and ad load. Competitors add pressure by periodically raising customer quality expectations. Broad product development and research can make competitors catch up faster, while targeted development for specific groups is harder to copy and lets competitors catch up more slowly. We show the computation of a customer’s perceived product quality and breakdown of each

5

(a) Maximize realism with

(b) Robust simulation through

granular simulation.

Example: Each customer is simulated individually with distinct behaviors. FREELANCERS

ENGINEERS

$

$

$

$

$

$

$

$

$

$

$

$

(c) Interconnected

mechanistic rules.

world dynamics.

Example: A price-vs-quality curve keeps customer behavior consistent.

Example: Customer segments influence each other and prevent exploitation of isolated dynamics

(d) Non-stationary environment. Example: Competitors shifts customer preferences and test agent’s capability to adapt Agent R&D finished quality +0.18 ↑

subscribe

Competitor moves

Engineers

quality

FINANCE

HEALTHCARE

$

$

$

$

$

$

$

$

$

$

$

$

Freelancer

quality +0.12 ↑

Customer expectation increases

reject $ price

Healthcare

$

Finance

Figure 5. Major design principles behind C E O - B E N C H ’s world mechanics and example designs that follow the principles.

factor in Equation 3. We describe full details of each factor in Appendix A.3, A.4, and A.6.  perc

Qi,t | {z }

quality perceived by customer i

=

mp |{z}

model-tier effect

+

  

q0 |{z}

+

initial quality

β r (ri,t − r0 ) | {z }

customer relationship

bshared |t {z }

+

dev improvement

+ β d log(αd + di,t /d0 ) − {z } | customer stickiness

group

bg,t | {z }

 − 

targeted dev improvement

β o ot |{z}

overload penalty

− β out 1{outaget } {z } | outage penalty

  U p,t − ηiads aeff − βU νU − i,t DU ui + | {z } {z } in-app ads penalty | open issues penalty β I Ii,t | {z }

quota saturation penalty

(3) Changing world imposes challenges. The world evolves over time through macroeconomic trends, interconnected reputation propagation, market saturation, demand surges, and competitor pressure. These factors affect acquisition, retention, and enterprise deal outcomes. The challenge is that the agent observes only partial and delayed evidence of these changes. It must infer hidden customer and market conditions from traces, choose actions whose effects arrive on different time scales, and revise its policy as the company and market move. Section 2.3 explains the interface design, and Section 2.2 describes the design principles that make the simulator mechanics realistic and challenging.

2.2 How We Make CEO-B E N C H Rigorous and Challenging We design CEO-B E N C H ’s world mechanics to be an expressive emulation of the real world, while remaining mechanistic so that success depends on genuine skills rather than exploiting brittle simulations. We describe seven core principles in our world mechanics design below and illustrate four examples in Fig. 5. Maximize realism with granular simulation. The simulator models 26 customer groups and individual customers within each group rather than only aggregate demand. Each customer has its own acquisition path, subscription state, price exposure, usage, satisfaction, and cancellation trajectory. Customers are also organized into diverse groups with different needs, budgets, price sensitivities, ad channel effectiveness, support expectations, and behavioral patterns. This granularity increases the complexity of world dynamics and widens the set of viable strategies. Robust simulation with mechanistic rules. The world emulates real business behavior while maintaining stable cause-and-effect relationships. Almost all simulator outcomes are generated by explicit mechanisms rather than by using an LLM as an opaque judge. For example, customers decide whether to subscribe by comparing product value against price through a microeconomics-motivated participation rule (Mussa and Rosen, 1978). This design aims to avoid failure modes in benchmarks such as Vending-Bench (Backlund and 6

Distribution

Example use in simulator

Motivation

Normal

R&D project quality gain

Captures uncertain payoff

Poisson

Daily new prospective customers for a group

Models rate-based counts

Bernoulli

Involuntary cancellation event

Models binary shocks

Uniform

Reputation damage noise

Adds bounded uncertainty

Log-normal

Competitor quality-jump magnitude

Models skewed positive shocks

Table 2. Stochastic mechanisms in C E O - B E N C H . The simulator uses a variety of stochastic variables to model realworld uncertainties.

Petersson, 2025a;b), where an LLM-simulated supplier can reward agent’s unrealistic verbal promises. Consistent simulation under stochasticity. While we inject stochasticity into world dynamics to emulate real-world noise, we maintain consistency across runs with independent random number generators for different simulator components. For example, under the same random seed, after calling the market research tool multiple times, the agent always discovers the same sequence of new market groups, independent of actions in other areas. Hidden information and indirect feedback. CEO-B E N C H tests whether agents can gather information in a partially observable world. The agent receives only information that a real start-up manager could plausibly observe: dashboards, database records, social-media posts, research reports, and negotiation history. It does not observe true customer satisfaction, latent willingness to pay, churn propensity, competitor schedules, or demand parameters. Instead, it must infer these hidden variables indirectly, for example, by gauging customer satisfaction and complaints through social media or detecting competitor moves by analyzing cancellation behavior. Interconnected world dynamics. We design the simulated world to make it difficult to isolate a single causal relationship and hill-climb on it. Every decision can influence many other parts of the market. For example, reputation propagates across related groups, so a quality failure in one enterprise group can spill into nearby groups and eventually affect consumer demand. Increasing satisfaction of influential customer groups can boost growth more effectively than ads. Delayed and uncertain consequences. Many actions have delayed and uncertain effects, forcing longhorizon decision making under uncertainty. Costs may appear immediately, while the corresponding revenue, retention, research, or reputation effects arrive weeks later. R&D projects have stochastic completion timelines and quality improvements, so investing more does not deterministically produce an immediate gain. Enterprise negotiations also unfold over stochastic delays, making it costly to wait too long but risky to overreact to any single turn. We show the types of distributions used and example usage in Table 2. Non-stationary environment. Agents must continually gather new information and adapt because the environment changes over the course of a simulation. Competitors place adaptive pressure on product quality. Customer behavior also drifts over time, with different groups shifting at different rates in price sensitivity and quality expectations. Macroeconomic trends add another changing background process, affecting willingness to pay and enterprise seat counts across expansions and contractions.

2.3 A Versatile Action Interface Between World and Agent We design a programmable tool interface, so agents can effectively manage granular action spaces and organize them into custom workflows. Composable action interface in Python. Terminal-based computer-use agents have become a general form factor across tasks (Anthropic, 2026; OpenAI, 2026; OpenCode, 2026; Pi Contributors, 2026). We make evaluating C E O - B E N C H easy with any of these agents by exposing the action surface to the agent via a Python package, novamind_api. An agent manages the company by calling functions in novamind_api in a Python script and executing the script in its terminal. This design maximizes flexibility for an agent to build 7

Database enables realistic analytics workflow

Granular action space widens possibilities

Composable tools allow sophisticated workflows

Example. Agent analyzes financials through SQL queries.

Example 1. Allocate ad spend by channel and customer segment.

Example. Agent connects database records directly to promotion decisions.

analytics.py

ad_spend.py

cash = query(""" SELECT SUM(amount) AS cash FROM ledger """)

marketing.set_targeted_ad_spend( targeted_spend={ "social_media": {"S1": 120}, "linkedin": {"S3": 100}, })

# recurring revenue from active subs mrr = query(""" SELECT SUM(seat_count * effective_price) AS mrr FROM subscriptions WHERE status = 'subscribed' """) print(f"cash={cash}") print(f"mrr ={mrr}") STDOUT

cash=−$33,970 mrr =$847,294/mo

retention.py

Example 2. Direct ops spend at specific at-risk customers. ops_spend.py at_risk = [ 4821, 5103, 5277, 5394, 5612, 5901, ]

rows = query(""" SELECT s.customer_id, s.effective_price FROM subscriptions s JOIN customers c USING(customer_id) WHERE c.group_id='S3' AND s.plan='B' AND s.effective_price > 79 """)["rows"] avg = sum(r["effective_price"] for r in rows) / len(rows) promos = {str(r["customer_id"]): round((r["effective_price"] - avg) * 0.4) for r in rows}

ops_spend.set_targeted_ops_spend( by_customer={ cid: 150 for cid in at_risk })

pricing.set_promotion(by_customer=promos) print(f"{len(promos)} subs, total {sum(promos.values()):,}/mo") STDOUT

134 subs, total $1,847/mo

Figure 6. Agents interact with C E O - B E N C H through a versatile Python interface. Left: We give the agent access to diverse business databases to test its information acquisition capability through a realistic data analytics workflow. Middle: We widen agents’ opportunity space by enabling them to take fine-grained actions. Right: This interface design allows the agent to compose tools into sophisticated custom workflows.

its own infrastructure on top of the API. In Fig. 6 (right), we show an example where, rather than calling a tool once per customer, an agent connects to the database via its custom data-driven promotion management system and applies promotion decisions efficiently at scale. Granular action spaces. We allow agents to act at fine granularity to create a rich space of strategic tradeoffs, failure modes, and opportunities for adaptation. Although the interface contains a finite set of tools, each tool accepts fine-grained structured arguments, so agents can compose a combinatorially large space of possible actions. In Fig. 6 (middle), we show examples where the agent allocates advertising spend by (ad channel, customer group) pair and decides operations spending on individual customers. Large-scale and realistic databases. We give the agent access to a 19-table operational database covering orders, contracts, subscriptions, the cash ledger, the social-media feed, configuration history, ad-channel attribution, and support tickets, among others. The schema mirrors what a real software company’s analytics stack would expose, testing the agent’s capability to gather information via an analytics workflow that resembles real-world software company operations. In Fig. 6 (left), we show an example where the agent analyzes its revenue through database queries. Social media. The agent can read a simulated public feed of customer complaints, competitor announcements, and macroeconomic trends. Agents can also reply and post on social media. Reactions to the agent’s posts on social media can also influence the rate of new customer acquisition. We test the agent’s capability to both perceive and act in a chaotic natural-language domain.

3 Experiments and Results In this section, we describe our experiments and their results. We then conduct both qualitative and quantitative analysis to compare behaviors across models. 8

Bankruptcy

Max final cash ($)

Max survival days

Mean survival days ± std

Turns /week

Best run API cost

Claude Opus 4.8 GPT-5.5 Claude Opus 4.7 Kimi K2.6 Claude Sonnet 4.6 GLM 5.1 Claude Haiku 4.5 Gemini 3 Flash DeepSeek V4 Pro Grok 4.20

0 2 0 1 2 3 3 3 3 3

27,776,973 21,297,707 389,959 98,050 69,766 0 0 0 0 0

500 500 500 500 500 324 231 226 176 37

500.0 ± 0.0 333.7 ± 229.7 500.0 ± 0.0 343.0 ± 110.0 282.3 ± 136.0 214.7 ± 91.1 144.7 ± 70.5 154.0 ± 37.0 114.3 ± 38.6 28.3 ± 8.5

10.8 34.7 14.6 30.5 13.3 51.5 23.1 18.5 19.3 8.2

$213.41 $200.49 $128.72 – $82.84 – $6.68 $2.98 – $0.75

Rule-based baseline Estimated final cash upper bound

– –

15,756,408.06 2,200,000,000

500 –

– –

– –

– –

Model

Table 3. Benchmark results summary. Most models fail to avoid bankruptcy, while Claude Opus 4.8 and GPT-5.5 finish above the initial $1,000,000 cash balance. The best model performance falls short of the estimated upper bound of attainable final cash by a large margin. CEO-B E N C H presents a challenging task for existing models.

3.1 Experimental Setup Models. We evaluate agents on the full 500-day CEO-B E N C H simulation. Each model is given $1M starting cash. We run three simulations for each model with random seed 42. We evaluate closed-weight models (GPT-5.5 xhigh, Claude Opus 4.8 max, Claude Opus 4.7 max, Claude Sonnet 4.6 max, Claude Haiku 4.5 thinking, Gemini 3 Flash Preview high, and Grok 4.20 Reasoning) and self-hosted open-weight models (DeepSeek-V4-Pro reasoning, GLM-5.1 reasoning, and Kimi-K2.6 reasoning). Harness. Terminal-based computer-use agents have become a general interface for automation: systems such as Claude Code, Codex, OpenCode, and Pi can perform diverse tasks and maintain memory by interacting with a terminal (Anthropic, 2026; OpenAI, 2026; OpenCode, 2026; Pi Contributors, 2026). We design C E O B E N C H to be compatible with any such agent. To align the harness across all models, we implement a minimal terminal agent interface: we give each agent a Linux working directory and tools including bash, read-file, and edit-file. In early runs, we found that open-source harnesses such as OpenCode and Pi (pimono) did not manage context reliably enough for 500-day episodes, so our harness refreshes context by clearing action history and only keeping system prompt and an agent-editable memory file in context at the start of each simulated week. Result selection for analysis. For results and analysis, we select the best run for each model as follows: (1) if at least one of the three runs avoids bankruptcy, we choose the run with maximum ending cash; (2) if all runs end in bankruptcy, we choose the run with the maximum number of simulation days before bankruptcy.

3.2 Results Overview Overall results. We show the best-run cash over time for each model in Fig. 2, and per-model trajectories across all three runs in Fig. 3. We also show additional details in Table 3. Most state-of-the-art models struggle to complete the simulation without bankruptcy. While five models (Claude Opus 4.8, GPT-5.5, Claude Opus 4.7, Kimi K2.6, and Claude Sonnet 4.6) end with positive cash on their best run, only Claude Opus 4.8 and GPT-5.5 finish above their $1M starting balance. This preliminary evaluation shows that Claude Opus 4.8 and GPT-5.5 demonstrate high-upside strategic behavior; Claude Opus 4.7 survives more conservatively; and most models fail to coordinate growth, quality, and cash flow. Rule-based baseline. We include a simple rule-based heuristic baseline that uses no language-model calls during policy execution: it fixes prices, quotas, and model tiers, concentrates acquisition and targeted development on a small set of customer groups, and adjusts capacity from recent usage. We conduct a preliminary grid search over this rule template and display the best strategy in Fig. 2 and Table 3. The heuristic achieves a significant positive cash balance of $15.76M, with Claude Opus 4.8 and GPT-5.5 exceeding it. We show full details of this baseline strategy’s design and configuration search in Appendix B. Benchmark is far from saturated. We estimate loosely the upper bound of achievable final cash be around 9

(a) Claude Opus 4.8 memo examples Day 7

Day 203

Day 266

Day 371

“I concentrated S1 ads on content marketing, dropped S3, raised Plan A quota, tilted dev to unlock S2, and added a small S2 ad probe.”

“Reallocating ads by measured ROI: SCALED D_S07 $12k→$30k/day, PULLED saturating S2 $100k→$85k/day, and nudged S1.”

“HARVEST PIVOT… conversion cratered, so CUT S2 + D_S07 ads, CUT D_S04 dev, TRIM S1 dev, and KEEP +EV engines.”

“No positive-EV action remains, so I’m coasting to liquidation with NO config changes (already minimal-cost: T0, model T1, zero ops/dev/ads).”

Day 112

Day 245

Day 287

“Mid-late game harvest+maintain. CUT D_S10 ads, REDIRECTED freed dev into new engine D_S06, and BUMPED ops to clear issues.”

“Cut expenses: capacity T6→T5 and global ops $40k→$28k… burn $88k→$58k/day while holding dev at $0 and S1 ads for a clean read.”

Day 497

“Light touch: reallocated dev — cut over-protected S1 targeted spend and bumped S2 to 5000/day to accelerate the growth-quality experiment.”

Day 7

Day 63

Day 147

Day 273

“I stopped broad ads and moved to a cash-conserving validation posture… kept S2/E3 targeted dev… left only a cheap S1 content/referral test.”

“I raised new-customer prices and quotas together, doubled S2 content spend to 16k/day… moved to tier4 capacity… started tier2 R&D…”

“I cut saturated S2 content ads to 35k/day, scaled proven D_S07 search/linkedin to 26k/day…”

“Churn spiked after competitor shocks… so I raised models to A3/B4/C5 as a temporary quality bridge…”

Day 21

Day 91

Day 203

Day 420

“I shifted to scaling S2 content at $1k/day, stopped weak S2 search/S1 paid ads/E3 content, kept E3 LinkedIn optionality…”

“I upgraded to capacity tier6… added D_S07 dev and promo, started tier6 R&D, and discovered/researched new groups.”

“I restored A2/B3/C4 quality, paused D_S07 paid acquisition, kept only positive S2 content ads, shifted ops toward D_S07 support…”

“I raised A/B tiers to A3/B4, restored S2 B promo s to 60k/day…”

“NO config changes… and NO actions… Just coast. This advance lands on d504, the FINAL day where the sim freezes.”

(b) GPT-5.5 memo examples

(c) Claude Opus 4.7 memo examples Day 77

Day 112

Day 210

Day 371

“DECISION: hibernate acquisition, invest quietly in quality. Actions: (1) killed all ads (save $4,200/wk), (2) reduced ops $150->$100/day…”

“HARVEST MODE CONTINUES -STATUS QUO. No config changes… Ads $0 (1.1% conv unprofitable)… Dev $0… No config changes.”

“HOLDING config unchanged (now 30+ weeks stable): prices A=$14/B=$49/C=$129, tiers 2/3/3, quotas 500/500/1500, capacity T0, ops $200/day, dev $0, zero ads.”

“HOLD & DIE strategy continues: T0 cap, 0 ops, 0 dev, 0 ads, no promos. Config unchanged Wk35→Wk53…”

Day 91

Day 147

Day 287

Day 490

“Continue HARVEST mode: keep all configs identical… No R&D (would gut cash, poor ROI given dying sub base). No ongoing promo…”

“HOLDING STATUS QUO - no config changes. Ops $200/day, dev $0, ads $0, capacity T0. Strategy working: book sticky, costs minimal.”

“STRATEGY: HOLD & DIE… All levers minimized: Capacity T0… Ops/Dev/Ads all $0… No R&D… Prices unchanged… No promos…”

“No changes: hold config… all levers offer zero upside with dead book. Projection Wk71 end (D497): $390,554 $595 = $389,959.”

Figure 7. Example memos written by Claude Opus 4.8 (top), GPT-5.5 (middle), and Claude Opus 4.7 (bottom) in their workspaces during the best trajectory of each model. Claude Opus 4.8 and GPT-5.5 actively explores and adjusts across diverse strategies, while Claude Opus 4.7 largely confines its decisions to a single strategic direction.

$2.2B. The estimation sums revenue from all 26 customer groups under maximum supportable pricing and subtracts the required costs for compute, capacity, development, operations, advertising, research, and acquisition. To obtain a conservative estimate, we further adjust downward for execution frictions, including issue-driven churn, enterprise negotiation friction, and acquisition delays. The resulting estimate remains far above the best observed model performance, indicating that C E O - B E N C H is far from saturated. We detail our estimation process in Appendix D. We release agent trajectories of all our experiments in an interactive trajectory viewer.

3.3 A Look into Agent Behaviors In this section, we take a preliminary exploration in agent behaviors and compare them across models. Strong models explore wider strategy space. In Fig. 7, we show example memos written by agents over time. GPT-5.5 and Claude Opus 4.8 adapts frequently as conditions change, trying a range of strategies such as scaling acquisition, adjusting model tiers, modifying promotions, and reallocating support or development spend. In contrast, Claude Opus 4.7 tends to respond to setbacks by repeatedly cutting spend and preserving

10

(a)

Claude Opus 4.8 · Simulating growth scenarios into the future · Day 77

(b)

D0=77; START_CASH=1269489 S1_TOUCHED0=159258+54467+987 RAW=[(0,19,15),(5,22,1830),(8,26,20562), ⋯ active cohorts ] def build_cohorts(): cohorts=[] ⋯ expand observed signup weeks into renewal cohorts return cohorts def simulate(cap=450000,conv=.72,first=.13,ongoing=.04): cohorts=build_cohorts(); touched=S1_TOUCHED0; cash=START_CASH REMAIN0=cap-S1_TOUCHED0 targets={84:'1wk',105:'4wk',161:'12wk',259:'26wk'}; results={} for step in range(1,183): day=D0+step; active=sum(c[1] for c in cohorts) pen=max(0.02,(cap-touched)/max(1,REMAIN0)) ad_spend=85000 if step<=7 else (120000 if step<=77 else 90000) leads=4.25*(ad_spend**0.71)*pen + 0.0132*active*pen new=leads*conv; touched+=leads; cohorts.append([26.0,new]) cash += new*15.0 - (ad_spend + leads*0.5 + active*0.15 + 58500) ⋯ apply re-bills/churn; record target days return results scen={'DISAST':(260000,.62,.25,.08),'PESS':(320000,.68,.18,.06), ⋯ scen 1wk M/k 4wk M/k DISAST .69/182 .07/152 PESS .99/202 .54/204 BASE 1.23/216 1.53/291 OPT 1.40/225 2.81/397

GPT-5.5 · Infer latent customer attributes from noisy negotiation outcomes · Day 133

rows=nm.query("""select et.customer_id,c.group_id,et.offer_json,s.st ⋯ 40+ lines suppressed def parse_prices(o): try: d=json.loads(o) except: return [] return [(x.get('plan'),x.get('price',x.get('price_per_seat',x.get( last={}; [last.setdefault(r['customer_id'],r) for r in rows] for gid in set(r['group_id'] for r in last.values()): lr=[r for r in last.values() if r['group_id']==gid] for plan in ['A','B','C']: by_status=collections.defaultdict(list) for r in lr: for p,v in parse_prices(r['offer_json']): if p==plan and v is not None: by_status[r['status']].append( if by_status: print(gid,plan,{s:(len(v),round(statistics.mean(v),2)) for s,v ⋯ 20+ lines suppressed Active enterprise accepted price distribution {'group_id':'E3','plan':'B','n':154,'seats':167595,'minp':1.69,'av {'group_id':'E3','plan':'C','n':3,'seats':3341,'minp':25.0,'avgp':

scen 12wk M/k 26wk M/k DISAST -4.58/131 -8.39/102 PESS -1.97/183 -1.91/158 BASE 2.79/273 8.99/248 OPT 10.74/462 29.96/433

Enterprise lost by latest agent offer and day groups summary GID D_E02 n 413 statuses Counter({'lost':403,'lead':10}) B all avg 40.19 min 39.0 max 49.0 by_status {'lost':(403,40.22,3 C all avg 61.15 min 59.0 max 79.0 by_status {'lost':(403,61.21,5 ⋯ 40+ lines suppressed

Figure 8. Example code files written by top-performing agents during their best trajectories. (a) The Claude Opus 4.8 agent runs its own simulation to forecast cash under different scenarios. (b) The GPT-5.5 agent infers latent enterprisecustomer price and quality preferences by mining noisy negotiation outcomes.

cash, which may help it survive until the final days but limits it from making positive profits. In Fig. 11, we find that Claude Opus 4.8 and GPT-5.5 distribute actions more evenly across tools than Claude Opus 4.7. Agents attain similar final cash through distinct strategies. Claude Opus 4.8 and GPT-5.5 in their best runs attain similar final cash balance. However, they attain this result via distinct strategies. In Fig. 9, we show their very different customer base change over time. Claude Opus 4.8 drops to zero customers mid-simulation, while GPT-5.5 sustains customers throughout the simulation, and the two agents focus on different customer groups. Fig. 7 also show that Claude Opus 4.8 decides to pivot to harvesting mode mid simulation and proceeds with expense cut and passive strategy maintaining. Sophisticated analytics by top-performing agents. In Fig. 8, we show example code files that top-performing agents write and execute. In (a), Claude Opus 4.8 constructs a cohort-based simulation to forecast future cash under different scenarios. In (b), GPT-5.5 mines negotiation history in the database to uncover hidden customer preferences. These examples demonstrate initial signs of sophisticated planning and information acquisition.

3.4 Measuring Drivers of Success and Failure While success in CEO-B E N C H requires multiple skills to work together, we conduct a preliminary analysis by isolating four skills and comparing them against agent performance. We compare quantitative measures of each skill for the top-performing models, against the average over all remaining models in Fig. 12, and explain each comparison further below.

11

Claude Opus 4.8

GPT-5.5

500k

Customer groups

500k

Individual 1

400k

300k

300k

200k

200k

100k

100k

sremotsuc evitcA

400k

Individual 2 Individual 3 Enterprise 1 Discoverable Individual 1 Discoverable Individual 2 Discoverable Individual 3 Discoverable Individual 4 Discoverable Individual 5 Discoverable Individual 6

0

0

100

200

300

400

500

0

0

Day

100

200

300

400

500

Day

Figure 9. Number of customers by customer group over time for the best runs of Claude Opus 4.8 and GPT-5.5. While Claude Opus 4.8 obtains more customers initially and drop to zero customer mid-simulation, GPT-5.5 sustains consistent customer base throughout. The two agents also focus on different customer groups. The two agents attain similar final cash balance via distinct strategy styles. Discoverable customer groups are initially hidden to agent and can only be discovered through paid market research. GPT-5.5 Opus 4.8 Opus 4.7 Kimi K2.6 Other Uncovering hidden information. In CEO-B E N C H , each pair of ad channel and customer group has 10% 11% 13% a different new-customer acquisition rate, emulat43% 44% 56% 57% ing real-world heterogeneity. The acquisition rate is 87% 89% 90% hidden from the agent, so the agent must uncover and use this information by analyzing customer Targeted dev spend Non-targeted dev spend acquisition history in databases. In Fig. 12a, we measure the average percentage of ad spending al- Figure 10. Targeted development spending breakdown. GPT-5.5 and Claude Opus 4.8 direct much larger shares of located to the best channel out of all ad spending development spending toward fine-grained group-specific for a customer group. We find that Claude Opus 4.8 improvements than most other models. and GPT-5.5 attain higher allocation efficiency than the remaining models. With five ad channels, the random-guessing baseline is 20%, and most models fall below that baseline.

Seeing into the future. In each simulated week, we ask the agent to submit a cash forecast four weeks into the future. Fig. 12b plots the percentage error between the submitted forecast and the realized cash balance four weeks later against the number of days before bankruptcy. We average data from the first four simulation weeks, when most models are still alive. We find that, on average, Claude Opus 4.8 has the lowest early forecast error, and the stronger models forecast with less error than the remaining models, demonstrating that they better understand the impact of their actions on the world. Speed of adaptation to environmental change. We measure adaptation speed with the time until the first occurrence of the word “competitor” in the agent’s workspace after the first competitor quality improvement. We show in Fig. 12c that, on average, Claude Opus 4.8, GPT-5.5, and Opus 4.7 detect environmental changes through indirect social media and database information faster than other models. Planning. We find that the stronger runs frequently anticipate different future scenarios and build corresponding solutions in their memos, with examples shown in Fig. 13. We show in Fig. 12d that Claude Opus 4.8 and GPT-5.5 use the word “if” more frequently than other models. Taking fine-grained actions. Our simulator allows agents to take actions at a fine-grained level. For example, agents can decide customer-group specific product development strategy. Proper analysis of customer group information and targeted development would result in advantages such as lower competitor pressure. Fig. 10 12

Claude Opus 4.8 1.06 1.03 0.72 0.72 0.65 0.63 0.61 0.43 0.39 0.36 0.5 1 1.5

get social posts list research projects get group insights set targeted ad spend set targeted dev spend get market overview set daily spend get cost info set capacity tier set lead promotion

0

GPT-5.5

Claude Opus 4.7 1.97 1.94 1.66 1.35 1.21 1.15 1.07 1.04 1.00 1.00 1 1.5 2

set targeted ops spend set targeted ad spend set promotion send enterprise deal post social media set targeted dev spend get group insights list research projects get cost info set capacity tier

Avg. uses per week

2

0

0.5

get social posts list research projects set promotion set daily spend get group insights get market overview set targeted ad spend set targeted dev spend get cost info post social media

Avg. uses per week

1.54 1.25

0.37 0.25 0.20 0.13 0.13 0.13 0.08 0.08 0 0.5

1

1.5

Avg. uses per week

2

Figure 11. Average per-week tool usage frequency for the best runs of Claude Opus 4.8, GPT-5.5, and Claude Opus 4.7 (top 10 tools per model). GPT-5.5 and Claude Opus 4.8 distribute actions more evenly across tools.

👍 Ads allocation efficiency 60% 40% 20%

43%

Random guess 33% 14%

0%

10%

👎

👎 Four-week forecast error

700% 600% 500% 400% 300% 200% 100% 0%

4

179%

Detect-competitor lag (weeks) 2

3

1

(a)

1

1

1

0

(b) Claude Opus 4.8

Claude Opus 4.7

4

2.63 2.74

2 0

(c) GPT 5.5

8.57 7.47

8 6

2 48% 8% 25%

10

Weekly avg. count of "if" in agent memo

(d) Other models

Figure 12. Better-performing models excel along four skill axes: (a) uncovering hidden ad-channel effectiveness and allocating spend to the best channel, (b) forecasting future cash, (c) reacting quickly to competitor events, and (d) planning more extensively. We show mean and standard deviation of each measurement in plots above.

shows the dollar-weighted split between targeted and non-targeted development spending. GPT-5.5 and Claude Opus 4.8 allocates almost 90% of development dollars to targeted improvements, compared with 10% for Kimi K2.6, and 43% for the remaining models. GPT-5.5 and Claude Opus 4.8 demonstrate a stronger tendency to take granular actions compared to other models.

4 Ablating Simulator Configurations We examine how simulation outcomes change when varying competitor and time-horizon configurations. We find that competitor difficulty provides an effective knob for tuning task difficulty, and our task remains challenging for existing models even over a short horizon.

4.1 Ablating Competitor Difficulty We ablate simulator difficulty by varying the competitor configuration. In C E O - B E N C H , the competitor raises customer expectations through both a preset stationary sequence and adaptive responses to agent actions. In the adaptive component, the competitor raises customer expectations by u · I, where u ∼ U [0.2, 0.5] and I is the agent’s cumulative quality improvement. We ablate simulator difficulty with the following settings: (1) stationary + adaptive competitor with u ∈ {0.1, 0.2, 0.3}; (2) stationary competitor only; (3) no competitor. In Fig. 14(a), we show that reducing competitor strength significantly reduces the difficulty of the task, and removing the competitor makes the task much easier. The ablation shows that competitor strength can be an effective knob for tuning task difficulty, and the non-stationary environment is a crucial component of what makes the task challenging.

13

capacity (a) Reversible downgrade

DAY 329

Scaled capacity tier7 -> tier6 because no enterprise seats and peak individual usage <80M/day; saves $329k/wk. If enterprise seats accept and usage/latency jump, go back to tier7 immediately.

(b) Hypothesis-test branching

(c) Customer group probing

DAY 336

DAY 63

Probing linkedin D_E09 to test whether enterprise threads arrive. If threads arrive, offer plan C or B near $80-$106 per seat; if no leads, bump probe or drop.

If the $40 180-day credits materially reduce churn, keep targeted renewalcredit strategy and prune expired cohorts weekly. If credits fail, avoid broad discounts and focus on quality.

GPT-5.5

(d) Price cut consequences

DAY 357

Cutting price has no downside when survival is 0%. If it flips renewals, keep the base and test elasticity; if it still fails, accept liquidation and coast.

Claude Opus 4.8

Figure 13. Examples of planning in GPT-5.5 and Claude Opus 4.8 memos. The agents anticipate scenarios and solutions with “if-then” contingencies. We show in Fig. 12(d) that these models anticipate more frequently than other models. $10M $1B

GPT-5.5

No competitor Stationary competitor

$10M $1M $100k

Adaptive competitor (u=0.1)

Cash (USD, log)

Cash (USD, log)

$100M

Adaptive competitor (u=0.2)

$10k $1k

Claude Sonnet 4.6

$1M

Claude Haiku 4.5 Kimi K2.6 GLM-5.1 DeepSeek V4 Pro

$100k

Gemini 3 Flash

$10k

Adaptive competitor (u=0.3)

$1k

Default config

0

100

200

300

400

500

Grok 4.20

0

10

20

30

40

50

Day

Day

Figure 14. Ablating simulator configurations. (a) Weaker or absent competitors make the task substantially easier. (b) Shortening the horizon to 50 days results in most models still unable to make profits.

4.2 Ablating Time-Horizon We examine whether agents behave differently when told to maximize cash balance over a shorter horizon. In this experiment, we change the simulation period to 50 days, one-tenth of the original simulation period. While the shortened horizon reduces challenges in long-term planning, Fig. 14(b) shows that only GPT-5.5 is still able to make a positive profit at the end. This analysis reveals that most models today remain weak in orchestrating decisions toward a short-term goal.

4.3 Ablating Agent Harness C E O - B E N C H can easily evaluate any agent harness. While we obtain most results with a custom minimal terminal-using agent harness, we ablate popular agent harnesses while keeping the underlying model fixed. For Claude Opus 4.7, we compare results with Claude Code (Anthropic, 2026), and for GPT-5.5, we compare results with Codex (OpenAI, 2026). In Fig. 15, we show that switching harnesses massively changes agent behaviors. Agents take significantly fewer actions when using Claude Code and Codex, resulting in inferior performance. While we cannot access full implementation details of these harnesses, we hypothesize that the difference results from software engineering-oriented system prompts of these harnesses.

14

$100M

Claude Opus 4.7

$100M

$10M

$10M

$1M

$1M

$100k

$100k

$10k

$10k

$1k

$1k

$100

0

100

☠️☠️ 200

300

400

Day Opus 4.7 + Claude Code

500

$100

GPT-5.5

8

Avg. turns per day 4.79

6 4 2.20

2 0

☠️ 100

Opus 4.7 + Our Harness

☠️200

300

400

Day GPT-5.5 + Codex

☠️ 500

1.27

0.32

0

GPT-5.5 + Our Harness

Figure 15. Cash trajectories and action frequency when ablating agent harnesses. We compare Claude Opus 4.7 and GPT-5.5 under our minimal terminal-using agent versus Claude Code and Codex respectively. Under Claude Code and Codex, the agents produce fewer actions per turn and achieve inferior performance.

5 Related Work Language model evaluations. Language-model evaluation has moved from static knowledge and reasoning (Hendrycks et al., 2021; Srivastava et al., 2023; Liang et al., 2023; Rein et al., 2024) toward realistic agentic task execution (Chen et al., 2021; Jimenez et al., 2024; Zhou et al., 2024; Xie et al., 2024b; Drouin et al., 2024; Trivedi et al., 2024; Yoran et al., 2024; Yao et al., 2025; Liu et al., 2024; Ma et al., 2024). Recent benchmarks have expanded evaluation scope to economically valuable deliverables in broad domains (Patil et al., 2025; Mialon et al., 2024; Chan et al., 2025; Starace et al., 2025; Miserendino et al., 2025; Patwardhan et al., 2025). However, their objectives usually terminate at a target state or one-shot deliverable. C E O - B E N C H instead asks whether agents can sustain progress toward a distant objective as earlier decisions continue to shape later states. Long-horizon agent evaluation. Memory and continual-learning benchmarks test models’ ability to retain information over time (Bai et al., 2024; Hsieh et al., 2024; Wu et al., 2025; Hu et al., 2026; He et al., 2026b; Laskin et al., 2023; Wang et al., 2023; Monea et al., 2024), but they are often limited to information retrieval or a single static task. Long-horizon benchmarks extend evaluation from isolated tasks to processes that unfold over time (Wu et al., 2024; Xie et al., 2024a; Xu et al., 2025; Chan et al., 2025; Starace et al., 2025; Wang et al., 2025; Luo et al., 2025; He et al., 2026a). Most recently, Vending-Bench asks agents to run a vending machine over many days, and AccountingBench asks agents to close monthly books from real software company data (Backlund and Petersson, 2025a;b; Penrose AI, 2025). However, they involve narrow operating problems, fewer coupled decisions, and largely stable or observable environments. CEO-B E N C H evaluates long-horizon agency in a broader operating setting where agents must coordinate pricing, growth, product, operations, communication, and enterprise sales under hidden state, noisy feedback, delayed consequences, and non-stationary market pressure in a consistent simulator. For example, we show in Appendix C that Vending-Bench allows models to accumulate successes relatively steadily, while our simulator requires an agent to make significant investments that only pay back much later, posing a stronger challenge to long-horizon planning.

6 Limitations and Conclusion 6.1 Limitations We make our best effort to approximate real-world startup operations and challenges in C E O - B E N C H . However, discrepancies can still exist between reality and the approximation. For example, since we have not found a reliable way to evaluate a model’s capability to propose qualitative changes to products, we simulate products using only a quality measure. In addition, to make each simulation run economically feasible, we limit the scope of possible actions and leave out aspects such as compliance, security, and fundraising.

15

6.2 Conclusion C E O - B E N C H shows a gap between existing models’ local tool competence and crucial sustained strategic skills: agents built on existing models can take plausible actions but fail when those actions must compound under delayed feedback, hidden state, and non-stationarity. To develop agents beyond isolated task executors, we need evaluations that ask whether they can organize evolving systems toward distant goals. CEO-B E N C H is one step toward that future: building agents and training models that do not merely answer requests, but help steer long-running organizations through uncertainty.

Acknowledgments We thank Modal for providing GPU resources for LLM inference. We thank Shuer Jiang, Boya Zeng, Sachin Konan, Taiming Lu, Linrong Cai, David Yin, Rahul Chalamala, Bryan Chiang, Luke Zeller, Yunyu Lin, Berkan Dokmeci, Bennett O’Brien, Ashank Tomar, and Ang Li for discussions and feedback. We thank Spencer Hong and The General Intelligence Company of New York for additional evaluations. KN acknowledges support from Schmidt Sciences.

References Anthropic. Claude code overview. https://code.claude.com/docs/en/overview, 2026. Axel Backlund and Lukas Petersson. Vending-bench: A benchmark for long-term coherence of autonomous agents. arXiv preprint arXiv:2502.15840, 2025a. Axel Backlund and Lukas Petersson. Vending-bench 2. https://andonlabs.com/evals/vending-bench-2, 2025b. Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. LongBench: A bilingual, multitask benchmark for long context understanding. In ACL, 2024. Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, Lilian Weng, and Aleksander Madry. MLE-bench: Evaluating machine learning agents on machine learning engineering. In ICLR, 2025. Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021. Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. WorkArena: How capable are web agents at solving common knowledge work tasks? In ICML, 2024. Muyu He, Adit Jain, Anand Kumar, Vincent Tu, Soumyadeep Bakshi, Sachin Patro, and Nazneen Rajani. YC-Bench: Benchmarking AI agents for long-term planning and consistent execution. arXiv preprint arXiv:2604.01212, 2026a. Zexue He, Yu Wang, Churan Zhi, Yuanzhe Hu, Tzu-Ping Chen, Lang Yin, Ze Chen, Tong Arthur Wu, Siru Ouyang, Zihan Wang, Jiaxin Pei, Julian McAuley, Yejin Choi, and Alex Pentland. MemoryArena: Benchmarking agent memory in interdependent multi-session agentic tasks. arXiv preprint arXiv:2602.16313, 2026b. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In ICLR, 2021.

16

Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. RULER: What’s the real context size of your long-context language models? In COLM, 2024. Yuanzhe Hu, Yu Wang, and Julian McAuley. Evaluating memory in LLM agents via incremental multi-turn interactions. In ICLR, 2026. Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. SWE-bench: Can language models resolve real-world GitHub issues? In ICLR, 2024. Daniel Kahneman and Amos Tversky. Prospect theory: An analysis of decision under risk. Econometrica, 47 (2):263–291, 1979. doi: 10.2307/1914185. Evan F. Koenig. Using the purchasing managers’ index to assess the economy’s strength and the likely direction of monetary policy. Federal Reserve Bank of Dallas Economic and Financial Policy Review, 1(6), 2002. URL https://fraser.stlouisfed.org/files/docs/publications/frbdalreview/frbdal_er02v01_n06_ a01.pdf. Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Hansen, Angelos Filos, Ethan Brooks, Maxime Gazeau, Himanshu Sahni, Satinder Singh, and Volodymyr Mnih. In-context reinforcement learning with algorithm distillation. In ICLR, 2023. Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. TMLR, 2023. Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, Shudan Zhang, Xiang Deng, Aohan Zeng, Zhengxiao Du, Chenhui Zhang, Sheng Shen, Tianjun Zhang, Yu Su, Huan Sun, Minlie Huang, Yuxiao Dong, and Jie Tang. AgentBench: Evaluating LLMs as agents. In ICLR, 2024. Haotian Luo, Huaisong Zhang, Xuelin Zhang, Haoyu Wang, Zeyu Qin, Wenjie Lu, Guozheng Ma, Haiying He, Yingsha Xie, Qiyang Zhou, Zixuan Hu, Hongze Mi, Yibo Wang, Naiqiang Tan, Hong Chen, Yi R. Fung, Chun Yuan, and Li Shen. UltraHorizon: Benchmarking agent capabilities in ultra long-horizon scenarios. arXiv preprint arXiv:2509.21766, 2025. Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He. AgentBoard: An analytical evaluation board of multi-turn LLM agents. In NeurIPS, 2024. James G. March. Exploration and exploitation in organizational learning. Organization Science, 2(1):71–87, 1991. doi: 10.1287/orsc.2.1.71. Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom. GAIA: A benchmark for general AI assistants. In ICLR, 2024. Samuel Miserendino, Michele Wang, Tejal Patwardhan, and Johannes Heidecke. SWE-Lancer: Can frontier LLMs earn $1 million from real-world freelance software engineering? In ICML, 2025. Giovanni Monea, Antoine Bosselut, Kianté Brantley, and Yoav Artzi. LLMs are in-context bandit reinforcement learners. arXiv preprint arXiv:2410.05362, 2024. Michael Mussa and Sherwin Rosen. Monopoly and product quality. Journal of Economic Theory, 18(2):301–317, 1978. Allen Newell and Herbert A. Simon. Human Problem Solving. Prentice-Hall, Englewood Cliffs, NJ, 1972. ISBN 0-13-445403-0.

17

Richard L. Oliver. A cognitive model of the antecedents and consequences of satisfaction decisions. Journal of Marketing Research, 17(4):460–469, 1980. doi: 10.1177/002224378001700405. OpenAI. Codex cli. https://developers.openai.com/codex/cli, 2026. OpenCode. Opencode: The open source ai coding agent. https://opencode.ai/, 2026. Shishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E. Gonzalez. The Berkeley Function Calling Leaderboard (BFCL): From tool use to agentic evaluation of large language models. In ICML, 2025. Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, Simón Posada Fishman, Marwan Aljubeh, Phoebe Thacker, Laurance Fauconnet, Natalie S. Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, Jerry Tworek, et al. GDPval: Evaluating AI model performance on real-world economically valuable tasks. arXiv preprint arXiv:2510.04374, 2025. Penrose AI. AccountingBench: Evaluating LLMs on real long-horizon business tasks. https://accounting.p enrose.com/, 2025. Pi Contributors. Pi documentation. https://pi.dev/docs/latest, 2026. David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google-proof Q&A benchmark. In COLM, 2024. Herbert A. Simon. A behavioral model of rational choice. The Quarterly Journal of Economics, 69(1):99–118, 1955. doi: 10.2307/1884852. Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. TMLR, 2023. Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, and Tejal Patwardhan. PaperBench: Evaluating AI’s ability to replicate AI research. In ICML, 2025. arXiv:2504.01848. Alexander Szimayer and Ross Maller. Testing for mean reversion in processes of Ornstein–Uhlenbeck type. Statistical Inference for Stochastic Processes, 7:95–113, 2004. doi: 10.1023/B:SISP.0000026032.80363.59. David J. Teece, Gary Pisano, and Amy Shuen. Dynamic capabilities and strategic management. Strategic Management Journal, 18(7):509–533, 1997. doi: 10.1002/(SICI)1097-0266(199708)18:7<509::AID-SMJ882>3.0. CO;2-Z. Harsh Trivedi, Tushar Khot, Mareike Hartmann, Ruskin Manku, Vinty Dong, Edward Li, Shashank Gupta, Ashish Sabharwal, and Niranjan Balasubramanian. AppWorld: A controllable world of apps and people for benchmarking interactive coding agents. In ACL, 2024. George E. Uhlenbeck and Leonard S. Ornstein. On the theory of the Brownian motion. Physical Review, 36(5): 823–841, 1930. doi: 10.1103/PhysRev.36.823. Weixuan Wang, Dongge Han, Daniel Madrigal Diaz, Jin Xu, Victor Rühle, and Saravan Rajmohan. OdysseyBench: Evaluating LLM agents on long-horizon complex office application workflows. arXiv preprint arXiv:2508.09124, 2025. Xiao Wang, Yuansen Zhang, Tianze Chen, Songyang Gao, Senjie Jin, Xianjun Yang, Zhiheng Xi, Rui Zheng, Yicheng Zou, Tao Gui, Qi Zhang, and Xuanjing Huang. TRACE: A comprehensive benchmark for continual learning in large language models. arXiv preprint arXiv:2310.06762, 2023.

18

Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. LongMemEval: Benchmarking chat assistants on long-term interactive memory. In ICLR, 2025. Yue Wu, Xuan Tang, Tom M. Mitchell, and Yuanzhi Li. SmartPlay: A benchmark for LLMs as intelligent agents. In ICLR, 2024. Jian Xie, Kai Zhang, Jiangjie Chen, Tinghui Zhu, Renze Lou, Yuandong Tian, Yanghua Xiao, and Yu Su. TravelPlanner: A benchmark for real-world planning with language agents. In ICML, 2024a. Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. In NeurIPS, 2024b. Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zora Z. Wang, Xuhui Zhou, Zhitong Guo, Murong Cao, et al. TheAgentCompany: Benchmarking LLM agents on consequential real world tasks. In NeurIPS, 2025. Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. τ-bench: A benchmark for tool-agentuser interaction in real-world domains. In ICLR, 2025. Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin, Ofir Press, and Jonathan Berant. AssistantBench: Can web agents solve realistic and time-consuming tasks? In EMNLP, 2024. Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. WebArena: A realistic web environment for building autonomous agents. In ICLR, 2024.

19

Appendix A Simulator Mechanics We describe full details of simulator mechanics in this section.

A.1 Agent Commands and Observable State Python action surface. The agent changes the company by importing functions from novamind_api and executing Python. The active command groups are: • pricing: set_prices, set_model_tiers, set_usage_quotas, and set_promotion. • marketing: set_daily_spend, set_targeted_ad_spend, and set_ads_strength. • marketing: set_lead_promotion and post_social_media. • analytics: set_targeted_ops_spend and set_targeted_dev_spend. • research: start_research_project and list_research_projects. • market: research_market, research_group, get_market_overview, and get_group_insights. • infrastructure: set_capacity_tier and get_cost_info. • enterprise: send_enterprise_deal and reject_enterprise_deal. • Time and data access: next_week advances the simulator, query reads the company database, and get_vars reports runtime variables. The interface is intentionally operational rather than omniscient: it exposes dashboards, database tables, social posts, and inbox messages, while leaving true preferences, satisfaction, competitor schedules, and hidden macro state latent.

A.2 Customers, Plans, and Participation Customer parameter sampling. Each customer belongs to a group g and receives private parameters such as maximum willingness to pay, quality floor, quality ceiling, usage demand, ad sensitivity, support sensitivity, and enterprise negotiation traits. For a generic customer parameter k,   min max θi,k,t = clip θ̃i,k + ∆market , θ , θ , θ̃i,k ∼ D g,k (µ g,k , σg,k ), (4) g,k g,k,t g,k where θi,k,t is customer i’s active value for parameter k, θ̃i,k is the sampled base value, D g,k is the configured group-level sampling distribution, µ g,k and σg,k are its location and spread parameters, ∆market is accumulated g,k,t min , θ max are clipping bounds. The distribution terms encode the fact that a market market drift, and θ g,k g,k segment has a typical profile, while the sampled base value gives each customer an idiosyncratic budget, tolerance, or usage pattern. The drift term represents changing market conditions, such as customers becoming more demanding over time, and the clipping bounds keep preferences in plausible ranges. This mechanism simulates real customer cohorts: people in the same segment resemble one another, but they are not interchangeable, and their preferences can move as the market changes.

Customer participation curve. Each customer i has maximum monthly willingness to pay ci , minimum quality floor qmin , quality ceiling qmax , and low-price and high-price slopes siL , siR . For offered effective price i i C, define normalized price x = C/ci , quality range ∆qi = qmax − qmin , and sigmoid σ (z) = (1 + e−z )−1 . With i i

20

L , θ R , θ max , ω , k , ψmin , ψmax , configurable curve coefficients θQ Q Q Q Q Q Q

    L R ψi (C ) = ωQ σ k Q siL ( x − θQ ) + (1 − ωQ ) σ k Q siR ( x − θQ ) ,   req max min . , ψQ + ∆qi clip ψi (C ), ψQ Qi (C ) = qmin i

(5)

Here ψi (C ) is the customer’s normalized required-quality score at price C; x = C/ci measures price relative L and θ R locate the lowto the customer’s willingness to pay; ∆qi is the customer’s personal quality range; θQ Q and high-price portions of the curve; ωQ blends the two portions; siL and siR make some customers more min , ψmax bound the score. The mechanic price-sensitive than others; k Q controls global curvature; and ψQ Q follows a participation-rule view of differentiated products: higher prices require higher perceived quality, and customers near their budget ceiling become more demanding (Mussa and Rosen, 1978). A non-enterprise max c and Qperc ≥ Qreq (C customer accepts plan p only if Ci,p,t ≤ θQ i i,p,t ). This simulates the real purchasing i,p,t i rule that customers do not compare price in isolation; they ask whether the product feels good enough for what they are being charged. Minimum accepted quality qi,min t (C)

1.0 0.9 0.8 0.7 0.6 0.5 0.4

Budget Budgetuser user Individual Individualpro pro SMB SMBteam team Enterprise Enterprise

0.3 0.2 0

50

100

150 200 250 300 Offered monthly price C (USD)

350

400

450

Figure 16. Example customer participation curves. Each curve represents the average participation behavior of a req customer group. Each curve maps the offered monthly price C to the minimum accepted quality Qi (C ) for a customer with different willingness to pay, quality floors, ceilings, and price-sensitivity slopes.

Plan choice and billing-period switching. Let P be the active plan set. For customer i, the simulator evaluates all plans and chooses the acceptable plan with the largest surplus: h i perc req perc req ∗ eff eff eff pi,t = arg max Qi,p,t − Qi (Ci,p,t ) , Ai,t = { p ∈ P : Ci,p,t ≤ ci , Qi,p,t ≥ Qi (Ci,p,t )}. (6) p∈P

∗ ∈ A ; if A is empty on a billing decision day, the customer cancels. P is The choice is valid only when pi,t i,t i,t eff is the price after active promotions, Qperc is perceived the set of plans currently offered by the company, Ci,p,t i,p,t quality for plan p, ci is the customer’s budget ceiling, and Ai,t is the acceptable-plan set. The surplus term measures how much perceived quality exceeds the customer’s minimum requirement at that effective price. This creates a simple operator intuition: customers do not maximize quality alone or price alone; they choose the plan that clears their personal price-quality bar with the best headroom. In reality, this corresponds to monthly subscription review: customers may upgrade, downgrade, stay, or churn depending on the menu of plans available at renewal time.

21

A.3 Product Quality, Usage, and Monetization Promotions and effective price. Promotions are additive across global, group, customer, and group-plan scopes. At billing time, h i global group group-plan eff Ci,p,t = Pp,t − Πt − Π g,t − Πcustomer − Π g,p,t − 1{firsti,t }Πlead . (7) g,t i,t +

global

Here Pp,t is the listed price for plan p, Πt

group

is a site-wide discount, Π g,t

targets a customer segment,

group-plan Πcustomer targets an individual customer, Π g,p,t targets a segment-plan pair, Πlead g,t is a first-bill lead i,t

promotion, firsti,t marks the first billing event for a newly acquired customer, and [·]+ floors price at zero. Promotions therefore act as dollar discounts rather than hidden quality boosts; they can help acquisition or retention but directly reduce revenue. The mechanism simulates couponing, contract discounts, and introductory offers, where revenue changes immediately even though the product itself has not improved. Usage and capacity. Each active subscriber has a daily usage draw ũi,t , plan quota U p,t , billing-period cumulative usage Ūi,t , and weekly usage multiplier Wt . The realized daily usage is  ui,t = round min ũi,t Wt , [U pi ,t − Ūi,t ]+ , Uttot = ∑ ui,t . (8) i

ui,t is the delivered usage units today, ũi,t is the customer’s latent demand, Wt is a weekly demand multiplier, U pi ,t is the quota on the customer’s active plan, Ūi,t is already-consumed usage in the billing period, and Uttot is total platform load. The positive-part term enforces the remaining plan quota, and rounding maps continuous demand into discrete usage units. The mechanic makes high-growth strategies stress infrastructure: more customers and higher quotas produce more usage, which can create overload if capacity is not upgraded. This simulates a real software business where usage is bursty, quotas cap consumption, and product-market success can turn into an infrastructure problem. ops

Service health. For capacity tier κt with capacity Kκt and operations spend xt , the overload level and outage probability are  ot =

 Uttot −1 , Kκ t +

  ops 0 P(outaget ) = max pmin , p exp (− x /χ ) (1 + βoout ot ) . ops out out t

(9)

ot is zero when capacity covers demand and positive when load exceeds capacity; Kκt is the capacity available ops under tier κt ; p0out is baseline outage risk; pmin is daily operations spend; χops out is the reliability floor; xt o controls diminishing returns to operations spend; and β out makes overload increase outage risk. Operations spending therefore buys reliability, while capacity tier choices buy headroom. This simulates the operational reality that SRE effort and cloud capacity reduce incidents, but overloaded systems remain fragile and no team can drive outage risk exactly to zero. Delivered quality. Delivered quality measures technical product quality before customer-specific perception effects:   group shared Qdel = q + b + b m p − β o ot − β out 1{outaget }, (10) 0 g,p,t t g,t group

where q0 is baseline product quality, btshared is shared quality from development and R&D, bg,t is targeted group quality for segment g, m p is the model-tier multiplier for plan p, ot is overload, and 1{outaget } is the outage indicator. The coefficients β o and β out translate infrastructure problems into quality loss. The productquality terms model the underlying capability of the service, while the negative terms model degraded delivery. This separates product investment from delivery reliability: a strong product can still feel bad when overloaded, just as real customers judge both feature quality and whether the service actually works when they need it.

22

Development and R&D. Daily development and targeted development add quality with diminishing returns: target

∆btshared,dev = β dev log(1 + xtdev /χdev ), target

xtdev is global development spend, x g,t

∆bg,t

target

= β target log(1 + x g,t

/χtarget ).

(11)

is targeted development spend for group g, ∆btshared,dev is the daily

target

shared-quality increment, ∆bg,t is the daily group-specific increment, β dev , β target convert spend into quality, and χdev , χtarget control diminishing returns. The logarithm makes the first dollars of engineering spend more productive than later dollars, reflecting coordination overhead and finite easy fixes. This simulates staffing and engineering allocation: basic improvements can be made quickly, but pushing quality further requires disproportionately more effort. R&D projects are larger delayed improvements: DR&D ∼ Drtime , j j

quality

GR&D ∼ Dr j j

,

btshared = btshared + ∆btshared,dev + +1

q

GR&D + ϵt . j

(12)

j: t=tdone j

j indexes a research project, r j is its tier, DR&D is completion delay, Drtime is the tier-specific time distribution, j j q

quality

GR&D is its quality gain, Dr j is the tier-specific gain distribution, tdone is the completion day, and ϵt j j is configured product-quality noise. The summation adds only projects that finish today, while daily development accumulates continuously. The design makes R&D a delayed investment: it can move the global quality frontier, but it does not instantly solve today’s churn risk. This simulates product roadmaps in which small engineering work compounds steadily, while larger research bets have uncertain delivery dates and payoffs. In-app ads. In-app ad strength is additive across global, group, and customer settings and then log-scaled:   global group log α a + κ a clip( at + a g,t + acustomer , a , a ) − log α a max min i,t aeff . (13) i,t = log(α a + κ a amax ) − log α a global

Here at

group

, a g,t

, and acustomer are configured ad strengths at company, segment, and customer scope; i,t

amin , amax are bounds; α a is a positive log offset; κ a controls how quickly raw ad strength saturates; and aeff i,t is the customer-visible ad load after clipping and saturation. The log scaling makes additional ad load less effective at the high end, matching the idea that an already ad-heavy product has limited extra monetization headroom. Ads create daily revenue ads eff Yi,t = ρads (14) i ai,t ni , where ρads is customer i’s ad-revenue sensitivity and ni is seats. The same effective ad load subtracts i from perceived quality. This creates a monetization tradeoff: ads are immediately lucrative but can reduce satisfaction and retention. The mechanism simulates ad-supported SaaS or freemium products, where more impressions produce revenue but also make the product feel noisier or less professional to some customers. Perceived quality. Perceived quality is the utility-relevant quality experienced by customer i after relationship, tenure, support, quota, and ad effects: perc

Qi,t

= Qdel g,p,t + β r (ri,t − r0 ) + β d log( αd + di,t /d0 ) − β I Ii,t   U p,t − ηiads aeff − βU νU − i,t , DU ui +

(15)

where ri,t is relationship score, r0 is neutral relationship, di,t is days subscribed, d0 is the tenure scale, αd is the tenure log offset, Ii,t is open-issue days, U p,t is plan quota, ui is sampled daily usage demand, DU converts daily demand to the quota period, and aeff i,t is effective ad load. Coefficients β r , β d , β I , β U control relationship, tenure, issue, and quota effects, while ηiads is the customer’s ad-quality sensitivity. The terms have direct 23

interpretations: good relationships and familiarity add tolerance; support delays, quota shortfalls, and ads subtract from experienced quality. This simulates the difference between engineering quality and customer experience: the same product can feel better to a long-tenured, well-supported customer and worse to a customer facing tickets, quotas, or intrusive ads.

A.4 Satisfaction, Retention, and Support Satisfaction. Instant satisfaction is quality surplus over the participation curve, perc

S̃i,t = Qi,t

req

− Qi (Ci,t ),

(16)

and stored satisfaction is an exponential moving average with configurable inertia λS : Si,t = λS Si,t−1 + (1 − λS )S̃i,t .

(17) req

Here S̃i,t is today’s surplus, Si,t is stored satisfaction, Ci,t is the effective current price, Qi (Ci,t ) is the perc quality the customer expects at that price, Qi,t is experienced quality, and λS is satisfaction inertia. This is an expectancy-disconfirmation design: customers are satisfied when experience exceeds the paid-price expectation and dissatisfied when it falls short (Oliver, 1980). The moving average means customers remember recent experience instead of resetting each day, so a bad outage or a good support recovery can affect future behavior for multiple periods. Downstream rules weight negative satisfaction more strongly, consistent with loss aversion (Kahneman and Tversky, 1979). The mechanism simulates customer sentiment as a memory-bearing state rather than a one-day reaction. Billing revenue and involuntary churn. On a billing day set by the billing period Dbill , subscription revenue is eff Ytsub = ∑ Ci,p n, (18) i ,t i i ∈Bt

eff is the effective price of customer i’s active plan, p is where Bt is the set of subscribers billed on day t, Ci,p i i ,t

the active plan, and ni is seats. Seat count multiplies revenue because an enterprise or team subscription pays for more users than an individual account. This simulates recurring subscription billing, where cash arrives in discrete renewal events rather than continuously every day. Before voluntary plan-choice churn, a group-level involuntary churn draw may occur:   invol invol invol invol µinvol = clip ϵ , 0, 1 , ϵinvol , σg ), Zi,t ∼ Bernoulli(µinvol (19) g,m g,m g,m ∼ N ( µ̄ g g,m ). invol and m indexes the billing period, ϵinvol g,m is the period-specific involuntary churn rate before clipping, µ̄ g invol invol σg are group-specific churn parameters, µ g,m is the clipped probability used for group g in period m, and invol Zi,t is the customer-level cancellation draw. This captures background churn such as procurement freezes, budget changes, company shutdowns, or stakeholder turnover that are not caused by the agent. Voluntary churn then follows the participation rule: if no plan clears the customer’s curve, the customer cancels; if a different plan gives higher acceptable surplus, the customer switches. The mechanism simulates the fact that some churn is controllable through product and pricing, while some churn is exogenous noise in the customer base.

Support issue generation and resolution. For a subscriber with no open issue, the issue probability is   P(issuei,t ) = clip ( p0 + pS (Sref − Si,t ) + pout 1{outaget })ni , pmin , pmax . (20) p0 is the base issue rate, pS converts low satisfaction into tickets, Sref is the reference satisfaction level, pout adds outage-driven issue risk, 1{outaget } activates that risk on outage days, ni is seats, and pmin , pmax bound the probability. More seats create more chances for someone to hit a problem, while poor satisfaction and 24

outages make support demand spike. This simulates customer-success queues where large accounts and unhappy users generate more tickets. Open issues are resolved by operations pools. For pool P with spend x P , group g members n g,P , and pool size | P|,   n g,P resolved Ng,t ∼ Poisson (bP + λ g x P ) . (21) | P| bP is the pool’s base resolution rate, λ g is group-specific operations efficiency, x P is spend assigned to support resolved is the number pool P, n g,P /| P| allocates capacity to group g according to its share of the pool, and Ng,t resolved. Global operations covers all open issues; targeted operations creates additional pools by group, plan, group-plan pair, or customer. Fast resolutions add relationship boosts, while unresolved issues increase open-issue days and decay relationship. The intuition is queue-based: more operations spend increases throughput, but only for customers covered by that pool. This simulates support staffing and escalation rules, where targeted customer-success effort can protect priority segments but cannot help customers outside the targeted pool.

A.5 Reputation, Social Media, and Acquisition Reputation impact. Each active customer contributes a daily reputation delta from satisfaction: ( ρ+ Si,t , Si,t ≥ S0 , rep δi,t = −ρ− |Si,t |, Si,t < S0 .

(22) rep

S0 is neutral satisfaction, ρ+ is the positive-reputation rate, ρ− is the negative-reputation rate, and δi,t is customer i’s daily reputation contribution. Negative satisfaction is allowed to have a different slope from positive satisfaction so that bad experiences can be more reputationally damaging than good experiences are helpful. This simulates word-of-mouth asymmetry in real markets, where angry customers often spread more salient feedback than mildly satisfied customers. For group g, rep

∑i∈ g δi,t ∆R g,t = logνN (max( Ng , Nmin )), max( Ng , Nmin )

R g,t+1 = clip( R g,t + ∆R g,t , Rmin , Rmax ).

(23)

Ng is the active subscriber count, Nmin is the small-sample normalizer, νN controls logarithmic scale, R g,t is group reputation, and Rmin , Rmax bound reputation. The averaging term prevents a single customer from dominating a large segment, while the logarithmic factor lets larger customer bases produce more visible aggregate reputation movement. This simulates customer reviews and public sentiment accumulating within a market segment. Cancellations add event damage   log (max( Ng , Nmin )) νN cancel Di,t = ηD ( β D + ξ t ) α D + χ D min(Si,t , S0 )2 , max( Ng , Nmin )

ξ t ∼ Uniform(ξ min , ξ max ),

(24) where ηD , β D , α D , χ D are configurable damage coefficients, min(Si,t , S0 )2 makes very negative satisfaction cancel is the reputation hit from cancellation. This makes visible especially costly, ξ t is event noise, and Di,t churn more damaging when the customer was very unhappy. The mechanism simulates public cancellations, angry posts, and negative references that can hurt a brand beyond the lost subscription revenue. Cross-group reputation spillovers. Discovered groups receive spillovers from related groups:   Rh,t+1 ← clip Rh,t+1 + ζ R Wg,h ∆R g,t , Rmin , Rmax .

(25)

Wg,h is the influence from group g to group h, ζ R scales spillover, ∆R g,t is the reputation change in the source group, and the clipping bounds keep the recipient group’s reputation in range. The mechanic makes 25

reputation networked: enterprise failures can affect nearby enterprise groups or adjacent market segments. This simulates reference networks, professional communities, and social adjacency, where the experience of one group can change expectations in a related group even before those customers use the product. Customer and agent social media. Customer social-media candidates are weighted by satisfaction extremity, negative satisfaction, recent satisfaction change, active service events, influencer status, and seat count:   post new 2 wi,t = ni ω inf ω · 1 + α | S |) · 1 + α [− S ] (26) ( − S i,t i,t + · (1 + α∆ | ∆Si,t |) · (1 + α E |Ei,t |) . g i,t post

new wi,t is the post-sampling weight, ni is seats, ω inf g is the group influence multiplier, ωi,t is the new-customer multiplier, αS weights satisfaction extremity, α− gives extra weight to negative satisfaction, α∆ weights recent satisfaction changes, α E weights active service events, ∆Si,t is the satisfaction change, and Ei,t is the set of active quality events such as outage, overload, issue, or quota frustration. The simulator samples up to Kpost posts per day from these candidates, so social media is a noisy but informative public signal rather than a complete survey. This simulates the selection bias of public feedback: large, influential, newly acquired, or upset customers are more likely to be heard than a random satisfied user. agent

Agent-authored posts are judged per discovered group. For group g, let e g,t ∈ [−1, 1] be the judged reaction score. The social multiplier entering acquisition is   agent A g,t = clip A0 + α A e g,t Vg,t , Amin , Amax , (27) agent

where e g,t is the judged reaction of group g to the agent’s post, Vg,t is exposure, A0 is neutral social effect, α A converts reaction-weighted exposure into lead impact, and Amin , Amax bound the multiplier. Public communication can therefore help or hurt growth depending on how each group reacts. This simulates product marketing and public relations: a message that resonates with one segment can accelerate acquisition, while a poorly received message can suppress demand. Daily new-customer generation. For discovered target group g, expected leads are ! xc,g,t Lc,g,t net λ g,t = R g,t Dg,t Ct Mg,t A g,t Zt ∑ + ∑ Nh,t Wh,g , xad c h

(28)

where R g,t is reputation, Dg,t is market availability, Ct is the calendar-cycle multiplier, Mg,t is the macro lead multiplier, A g,t is the agent-social multiplier, Zt is the active demand-surge multiplier, xc,g,t is channel spend, net is the referral matrix. The Lc,g,t is leads per reference ad spend xad , Nh,t is subscribers in group h, and Wh,g calendar-cycle multiplier makes demand oscillate through recurring seasonal cycles, so otherwise identical ad spend can perform better or worse depending on timing. The macro multiplier captures broad economic expansion or contraction; the social and surge multipliers capture communication effects and temporary external demand spikes; and the referral term captures word of mouth from existing subscribers. Market saturation is " !νD # Ng,t Dg,t = D0 − , n g,t ∼ Poisson([λ g,t ]+ ). (29) capg,0 (αcap + γg t/Ycap ) +

D0 is baseline availability, Ng,t is the current number of customers in group g, capg,0 is initial market capacity, αcap is the baseline capacity scale, γg is group capacity growth, Ycap is the time scale for capacity growth, νD controls saturation curvature, λ g,t is expected leads, and n g,t is realized leads. The positive-part operator prevents negative availability, and the Poisson draw converts the expected funnel volume into noisy realized leads. This combines paid acquisition, word of mouth, reputation, macro conditions, seasonal demand, temporary shocks, and finite market size. The mechanism simulates a real go-to-market funnel where the same budget can yield different outcomes depending on brand, timing, segment saturation, and randomness. 26

Demand surges. External demand surges are temporary acquisition shocks. For each active surge s, Zt = ∏ zs ,

surge

zs ∼ Ds

,

t ∈ [tstart , tend s s ).

(30)

s∈St

surge

St is the set of active surges, zs is the surge lead multiplier sampled from surge distribution Ds , and end are its active days. The product over active surges allows multiple external events to stack. Surges tstart , t s s create temporary windows where growth is easier, but the agent must still have pricing, quality, and capacity to retain the acquired customers. This simulates events such as press attention, industry shifts, or sudden demand spikes that increase inbound interest without guaranteeing durable revenue.

A.6 Market Discovery and Non-Stationarity Market and group research. Market research reveals new customer groups, while group research increases the information level for a known group after a delay: target ℓinfo = max(ℓinfo ), g,t , ℓ g,t+ Dresearch g,ℓ

research Dresearch ∼ D g, . g,ℓ ℓ

(31)

target is the requested level, Dresearch is the configured ℓinfo g,t is the agent-visible information level for group g, ℓ g,ℓ research is the delay distribution for that group and research depth. The max operator completion delay, and D g, ℓ means research can raise the information level but cannot erase already acquired knowledge. Results snapshot current market conditions when the research completes. The purpose is to make information acquisition an operational choice with time cost rather than a free static table. This simulates customer discovery, analyst work, and market research projects that improve visibility only after a delay and may already be slightly stale when delivered. end Competitor events. Competitor events are disabled before tstart comp and after tcomp . Let τ be the last event day. The mean interval is ( switch , ¯ t = m∆ ∆, t < t∆ ∆ (32) ∆, t ≥ tswitch , ∆

with a separately configured minimum interval ∆min . If t − τ < ∆min , no event occurs; otherwise ¯ t ). Et ∼ Bernoulli(1/∆

(33)

When an event occurs, a base boost is sampled and scaled over the run: sample Bt = clip(LogNormal(µ B , σB ), Bmin , Bmax )

αB + ρB

t − tstart B

!

ramp

TB

.

(34)

end ¯ tstart comp and tcomp bound the active competitor window, τ is the last event day, ∆t is the current mean interval between competitor events, Et is the event indicator, m∆ slows early events, ∆ is the baseline interval, tswitch ∆ is the interval switch day, and ∆min prevents events from arriving too close together. µ B , σB parameterize ramp event size, Bmin , Bmax bound it, tstart is the boost-ramp start day, and α B , ρ B , TB control magnitude scaling B over the run. Together these terms simulate rival launches that are not perfectly periodic, but become more serious as the market matures. The simulator also models adaptive competitor catch-up to the agent’s global unreleased global development and R&D gains. Let Ht be the unreleased shared-quality bank. With global

ut

global

global

∼ Uniform(umin , umax ),   sample global global Bt = max Bt , ut Ht ,

global

Ht+

global

= Ht

27

global

− 1{ u t

global

Ht

sample

> Bt

global

} ut

global

Ht

. (35)

sample

Thus Bt is the applied competitor quality boost, Bt global

global

is the exogenous shock, ut

is the random

catch-up fraction, and Ht+ is the remaining shared-quality bank after the event. The event boost is at least the stationary sampled shock but can be larger when the agent has accumulated large unreleased global improvements. If the adaptive term wins, the competitor consumes that fraction of the bank. This simulates competitors copying or matching broadly visible product improvements: large general advances can invite stronger competitive responses than small incremental changes. Targeted development has a parallel group-specific adaptive component: target

Dg,t target

where Hg,t

target

= v g,t Hg,t

,

v g,t ∼ Uniform(vmin , vmax ),

target

target

Hg,t+ = Hg,t

target

− Dg,t

,

(36)

is the unreleased targeted-development bank, v g,t is the group catch-up fraction sampled target

target

between vmin and vmax , Dg,t is the targeted drift shock, and Hg,t+ is the remaining targeted bank after catch-up. The applied event shifts all customers’ curves upward by adding Bt to global quality expectation, target κ g Bt to group g’s expectation, and Dg,t to group g’s expectation. Here κ g is group competitor reactivity. The intuition is that competitors are partly exogenous and partly responsive: broad R&D attracts broad catch-up, while targeted work is harder for competitors to fully copy but can still leak into group expectations. This simulates real competitive pressure where segment-specific improvements can create more durable advantage than broad, easily observed product gains. Macroeconomic schedule. The hidden macro state is an Ornstein–Uhlenbeck process around a sinusoidal PMI cycle (Uhlenbeck and Ornstein, 1930; Szimayer and Maller, 2004; Koenig, 2002): µt = µPMI + APMI sin(2πt/PPMI + ϕ) , PMIt+1 = clip(PMIt + ηPMI (µt − PMIt ) + σPMI ϵt , PMImin , PMImax ) .

(37)

with ϵt ∼ N (µϵ , σϵ2 ). µt is the current cycle target, µPMI is the baseline PMI, APMI is cycle amplitude, PPMI is cycle period, ϕ is phase, ηPMI is mean-reversion speed, σPMI is shock scale, and PMImin , PMImax bound the index. Intuitively, demand oscillates through a PMI-like expansion-and-contraction cycle, while the OU update makes temporary surprises fade back toward the current cycle phase. This simulates macroeconomic background conditions that drift gradually but also contain noise. For each customer group and macrosensitive dimension d,   PMIt − PMI0 Mg,d,t = max Mmin , M0 + β g,d . (38) PMI0 Mg,d,t is the macro multiplier, M0 is neutral, Mmin is its floor, β g,d is group sensitivity, and PMI0 is the neutral PMI reference. The dimension d can represent macro-sensitive quantities such as lead flow, willingness to pay, or enterprise deal velocity, and different groups react with different sensitivities. The simulator uses real-time hidden PMI internally, while the agent sees only period averages published every Dpublish days with delay Ddelay . This simulates management under delayed economic indicators: the world has already moved before the agent receives clean macro data.

A.7 Enterprise Sales and Negotiation Offer evaluation. Enterprise customers evaluate up to Koffer offered (plan, price) options and choose the one with highest perc req Sioffer = Qi,p,t − Qi (C ), (39) perc

where Koffer is the maximum number of options the agent can present, Qi,p,t is the perceived quality of the req

offered plan, Qi (C ) is the required quality at price C, and Sioffer is offer surplus. Positive surplus means the proposed plan clears the customer’s price-quality bar. If the best offer is positive, the customer accepts.

28

This simulates enterprise procurement evaluating a small menu of contract options rather than passively accepting a list price. Otherwise, the simulator computes req

Cimax ( Q) = max{C ≤ ci : Qi (C ) ≤ Q}

(40)

where Cimax ( Q) is the highest price the customer would accept at perceived quality Q, ci is the customer’s budget ceiling, and the max operator searches for the price that is still justified by the offered quality. This is the reservation price implied by the same participation curve used for self-serve customers. This simulates a procurement ceiling: the customer may negotiate, but there is a maximum contract value that the perceived product quality can support. With sampled initial factor f i and configured counter-offer decay γα , counter Ci,r = Cimax − γαr (Cimax − f i Cimax ) ,

(41)

counter is the counter-offer on turn r, f is the initial counter-offer fraction of the customer’s maximum where Ci,r i acceptable price, and γα controls how quickly later counter-offers approach that maximum. Early counteroffers start below true willingness to pay; later turns move toward the customer’s maximum acceptable price. If the configured maximum turn count is exceeded, the customer stops responding. This simulates negotiation anchoring: buyers may reveal willingness to pay gradually rather than immediately offering their ceiling. Reply delays are also stochastic:   reply reply ¯ di,r ∼ Di di /Mg,deal,t , σid , (42) reply

where di,r

is the delay before the customer responds, d¯i and σid are customer response-time parameters,

reply

Di is the customer-specific reply-delay distribution, and Mg,deal,t is the macro deal-velocity multiplier. Stronger macro conditions shorten expected delay by increasing the denominator, while slow conditions stretch sales cycles. This makes enterprise sales slower and less certain than self-serve conversion while still following the same underlying price-quality logic. The mechanism simulates procurement latency, stakeholder review, and macro-sensitive sales velocity.

A.8 Costs and Cash Flow Daily costs. Daily operating cost combines fixed infrastructure, variable compute, operations, development, advertising, targeted actions, lead acquisition, and research charges: capacity

Ktcost = Kκt

+ ∑ χp

usage

p

ops

use U p,t + xt

target-ops

+ xtdev + Xt

+ ∑ x g,t

target-dev

g

ads + ∑ xc,g,t c,g

(43)

group project + Ntlead clead + Ktmarket + Kt + Kt capacity

Kκ t

usage

is the fixed cost of capacity tier κt , χ p

ops

use is usage on that plan, x is per-usage cost for plan p, U p,t t target-ops

and xtdev are global operations and development spend, Xt

target-dev

is targeted support spend, ∑ g x g,t

ads is acquisition advertising spend across channels is targeted development spend across groups, ∑c,g xc,g,t

and groups, Ntlead clead is the per-lead acquisition charge, and the three K terms are market research, group research, and research-project charges paid that day. This cost equation simulates the operating budget of a software company, where fixed infrastructure, variable compute, staffing, marketing, and research all consume cash through different channels. The cash update is ads Bt+1 = Bt + Ytsub + ∑ Yi,t − Ktcost ,

(44)

i

ads is in-product ad revenue, and K cost is total where Bt is company cash, Ytsub is subscription revenue, ∑i Yi,t t operating cost. This is the main strategic coupling in the simulator: growth, quality, and reliability can create future revenue, but they consume cash immediately. The mechanism simulates startup runway management, where the agent must decide when to burn cash for future growth and when to preserve liquidity to avoid bankruptcy.

29

B Rule-Based Baseline Strategy and Configuration Search We use the rule-based baseline as a non-LLM point of comparison for the agent results in Section 3.2. The baseline is intentionally simple: it commits to a fixed pricing book, fixed product-quality and advertising spend levels, and a fixed customer targeting rule. It does not use market research, enterprise negotiation, social media analysis, promotions, or language-model calls. At the start of a run, the policy sets prices, model tiers, and usage quotas from one price book. During each simulated week, it selects target customer groups from its target rule. If cash is above a configured floor, it applies a small global development spend of $200/day, a fixed targeted development spend for each selected group, and, starting on day 20, a fixed targeted advertising spend for each selected group. Advertising is placed on the highest-yield channel for the selected group under the simulator’s fixed channel table. Operations spend is set to max(100, 0.05nt ) dollars/day, where nt is the current active subscriber count. The policy adjusts capacity by at most one tier per week toward the cheapest tier whose capacity covers recent average usage at 80% utilization. This simulates a simple operating playbook: spend proportionally to current scale, focus on a small target market, and avoid complex adaptation or hidden-state inference. The configuration search is the Cartesian product of the options in Table 4, giving 24 configurations. Each configuration is evaluated for the full 500-day simulation with seed 42. The best configuration uses the mid price book, targets S1 only, uses the heavy spend package, and has a $100K cash floor; the $300K cash floor gives the same result for this seed. The traced replay of this configuration ends with $15.76M in cash. Search dimension

Options

Price book

cheap: A=$8/T1/50K tokens, B=$18/T2/200K tokens, C=$40/T3/1M tokens; mid: A=$12/T1/60K tokens, B=$25/T2/250K tokens, C=$55/T4/1.2M tokens. S1 only: target S1 throughout the run; S1 then S3: target S1 initially, then target both S1 and S3 from day 30 onward. light: S1 dev/ad=$2K/$250 per day, S3 dev/ad=$1.5K/$600 per day; medium: S1 dev/ad=$4K/$500, S3 dev/ad=$3K/$1.2K; heavy: S1 dev/ad=$6K/$750, S3 dev/ad=$4.5K/$1.8K. Stop optional global development, targeted development, and advertising when cash is at or below either $100K or $300K. Advertising begins on day 20; capacity targets 80% utilization; market discovery and enterprise actions are disabled; default competitor events remain active.

Target rule Spend package

Cash floor Fixed choices

Table 4. Search space for the rule-based baseline. The baseline search varies two price books, two target rules, three spend packages, and two cash floors, for 24 total configurations.

C Comparison to Vending-Bench 2 Curves Vending-Bench 2 reports the average GPT-5.5 balance over five 365-day runs, with the leaderboard ending near $7.5K from a $500 starting balance (Backlund and Petersson, 2025b). Its curve is comparatively steady: once the agent finds workable suppliers and products, gains can accumulate through repeated restocking and sales. In contrast, the best GPT-5.5 CEO-B E N C H run begins with $1M, falls to roughly $430K by day 30 as it funds acquisition, development, capacity, and operations, and only later compounds to $21.3M by the end of the run. This delayed-payoff structure makes local cash preservation an unreliable proxy for success and requires the agent to keep investing through long intervals where the benefits of its actions have not yet fully materialized.

30

CEO-Bench GPT-5.5 Vending-Bench GPT-5.5

$7.5k

e c n al a b h s a c h c n e B - O E C

$16.2M

$5.7k

$11.1M

$4k

$6.1M

$2.2k

$1M

$500

Day 0

ecnalab hsac hcneB-gnidneV

$21.3M

Maximum simulation period

Figure 17. GPT-5.5 cash trajectories in Vending-Bench 2 and C E O - B E N C H . Vending-Bench 2 reports an average trajectory over five 365-day runs, where cash grows relatively steadily from a $500 starting balance. The best C E O B E N C H run starts at $1M, draws down early to fund growth and product investment, and only later compounds to a much higher ending balance, illustrating the delayed-payoff structure of our simulator.

D Upper Bound Final-Cash Estimate We use the estimated maximum final cash only to calibrate remaining headroom in C E O - B E N C H . It is not a mathematical proof of optimality. The estimate has two stages. We first compute a pre-friction accounting subtotal by summing revenue from all customer groups under maximum supportable pricing and subtracting the costs required by the simulator mechanics. We then apply a friction adjustment for issuedriven churn, enterprise negotiation friction, and acquisition delay. The reported estimate is the adjusted value, approximately $2.2B. Objective. Let G denote the 26 customer groups and let B = {30, 60, . . . , 480} denote the billing days. For each group g, let Ng be the maximum attainable customer count or enterprise seat count, p g be the maximum supportable price, and r g (t) be the retention probability at billing day t. These terms define an optimistic full-market revenue calculation: how many customers can be reached, what price they can support, and how likely they are to remain active at each renewal. The pre-friction revenue estimate is R = ∑ Ng p g ∑ r g (t).

(45)

t∈ B

g∈ G

Here R is total subscription revenue before execution frictions. The inner sum adds the retained billing opportunities for a group over time, and the outer sum aggregates across all customer segments. This simulates the revenue side of a best-case operating plan in which acquisition succeeds and retention is governed by the modeled quality checks. The pre-friction final-cash subtotal is then Cpre = C0 + R − K,

(46)

where Cpre is final cash before friction adjustment, C0 = $1M is the initial cash balance, and K is the sum of all modeled simulator costs. This subtotal simulates accounting profit before execution slippage: revenue plus starting cash minus the cost of capacity, compute, development, operations, acquisition, and research. The reported estimate applies a friction factor F: Creported = FCpre .

(47)

where Creported is the final headroom estimate after friction and F is the multiplicative discount applied to the pre-friction subtotal. We use F = 0.49 to make the estimate conservative with respect to issue-driven churn, enterprise negotiation friction, and acquisition delay. This simulates the gap between a clean accounting upper bound and a real execution path, where customers do not all arrive instantly, enterprise deals take negotiation, and support issues can erode retention. 31

Parameter choices. The accounting assumes all 26 groups are retained through all billing cycles before friction adjustments. The selected configuration uses T3 inference, T7 capacity, $40K/day in global development, targeted development for every group, all ten R&D tiers started in parallel on day 1, and operations spending of $500/day plus $0.001 per active subscriber. These choices were selected from a small grid over targeteddevelopment slack, model tier, global-development spend, R&D schedule, capacity tier, and operations spend. The most sensitive parameter is targeted development: a 1–2× multiplier did not reliably retain all groups late in the run, while a 3× multiplier maintained full-cycle retention in the tested configurations and a 4× multiplier added cost without improving retention. Quality and retention check. For each candidate configuration, we compute the shared-quality, competitordrift, and R&D-bonus trajectories over the operating horizon. For group g, targeted-development spend is first sized to cover the late-run quality threshold,     ∆g (0) Tg = 5000 exp −1 , (48) 0.7A g · 0.0225 (0)

where Tg is the one-pass targeted-development spend estimate for group g, ∆ g is the required targeted quality bonus after accounting for shared quality and R&D effects, A g is the number of active development days, 5000 is the spend scale from the simulator’s targeted-development rule, 0.7 is the targeted-development conversion coefficient used in this sizing approximation, and 0.0225 is the quality-gain coefficient. The exponential is the inverse of the simulator’s log diminishing-returns function: larger required quality gaps (0)

require disproportionately more spend. We then use Tg = 3Tg to account for group-level drift feedback that the one-pass sizing rule does not capture. This simulates planning with safety margin, where an operator overbudgets targeted product work because competitors and market drift can raise the bar during execution. At each billing day, retention is computed from the delivered-quality cushion,   m[q0 + qshared (t) + bg (t)] − d g (t) − ℓ(t) − µ g r g (t) = Φ , (49) σg where r g (t) is group retention probability at billing day t, Φ is the standard normal cumulative distribution function, m = 0.90 is the T3 quality multiplier, q0 is baseline quality, qshared (t) is shared quality from global development and R&D, bg (t) is the targeted group bonus, d g (t) is fixed and competitor-induced group drift, ℓ(t) is the capacity-overload penalty, and (µ g , σg ) parameterize the group’s quality threshold distribution. The numerator is the quality cushion above the group’s threshold, and dividing by σg converts that cushion into a probability under the assumed threshold distribution. Under the selected configuration, the computed retention check gives r g (t) ≈ 1 for all 26 groups and all 16 billing cycles. Full retention implies approximately 521M usage units per day; with T7 capacity, the resulting overload penalty is included in ℓ(t). This simulates a stress test for whether product quality, targeted improvements, and capacity are strong enough to keep every segment through repeated renewals. Revenue and costs. The pre-friction calculation gives $6.69B in subscription revenue: $1.93B from individual groups and $4.76B from enterprise groups. We subtract every cost category used by the simulator rather than only direct infrastructure costs. Table 5 shows the resulting cost accounting. This paragraph simulates the full-company accounting view of the upper bound: high revenue is meaningful only after paying for the compute, capacity, product work, support, advertising, and information acquisition required to keep that revenue. Interpretation. The final row of Table 5 is the estimate used in the main text and in Table 3. It should be read as an approximate headroom calculation, not as a demonstrated executable strategy. The pre-friction accounting assumes full conversion, full enterprise close rates, rapid acquisition, maximum supportable prices, mean R&D timing and quality effects, and no additional losses from unresolved customer issues, discounts, reputational effects, or cash-flow constraints beyond the cost categories above. The friction factor makes the estimate conservative by reducing this subtotal to roughly $2.2B. This remains far above the best observed agent performance and supports the conclusion that C E O - B E N C H is far from saturated. 32

Quantity

Amount ($M)

Subscription revenue Initial cash

6,690.0 1.0

Compute Capacity Global development Targeted development R&D projects Advertising Lead acquisition Operations Market research Group research

1,554.0 37.3 19.9 423.6 9.2 114.8 0.9 2.5 1.5 1.0

Total modeled costs Pre-friction ending cash subtotal Friction adjustment factor Estimated final cash upper bound

2,164.7 4,526.3 ×0.49 2,200.0

Table 5. Accounting for the estimated final cash upper bound of $2.2B after adjustments for execution frictions.

33

Related documents

Record · ID 287194 · SHA-256 184236a1771f8a60
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.