Agent Fleets Are a Budget Line, Not a Feature Flag

Published September 10, 2026 · Updated September 10, 2026By ABD Legacy LLC
Agent fleet costFan-out decisionVerification budgetAgency procurement

Quick answer: fan out only when subtasks are independent, verifiable and parallelisable — all three. The only published fleet-scale anchor is OpenAI’s 2026 run: 10,000 agents × 88 hours = 880,000 agent-hours, 2.7 million messages and approximately 130 billion output tokens, with verification the line item no quote includes.

“Add more agents” is the easiest sentence in an AI pitch and the hardest one to invoice. On September 8–9, 2026, OpenAI said an internal system put on the order of 10,000 concurrent agents to work on the Navier–Stokes existence and smoothness problem and reached a resolution in about 88 hours, exchanging 2.7 million messages and using approximately 130 billion output tokens. Whatever you think of the claim, that is the most concretely metered agent fleet anyone has published — and it is the first one you can price a client project against.

This result is disputed and unreproducible publicly. It ran on an internal, unreleased model, no independent verification exists as of this writing, and the Clay Mathematics Institute has accepted nothing. Nothing below asserts that the proof is correct or accepted; the published workload is the only part of the run this page treats as a fact.

Agency buyers are being sold fleets with no unit of account. The pitch’s unit is agents. The bill’s unit is agent-hours and messages — and, if the work is any good, a second line item nobody quotes: verification. This page is the decision guide we use: when to fan out at all, how to put the fan-out in agent-hours before approving it, and how to make a vendor quote the verification half instead of leaving it as the part of the bill that arrives last and is argued about longest.

Agent fleet cost: what actually gets metered

An AI agent fleet’s cost is set by agents multiplied by hours (agent-hours), not by the agent count in the pitch. The 2026 OpenAI run is the anchor: 10,000 agents for 88 hours is 880,000 agent-hours, and the two rates it implies are 147,727 (about 147.7K) output tokens and 3.07 messages per agent-hour — derived from the published totals, not OpenAI-published unit rates.

Here is the same run written as a pitch and as a meter. Every figure in the right-hand column is OpenAI’s own published claim, quoted from its post: the money language in the left column belongs to reporters and a third-party estimator, not to OpenAI, whose post contains no cost language at all.

What you get pitchedWhat the published run actually metered
“Add more agents”on the order of 10,000 concurrent agents — at the final stage of one problem, not an exact headcount
“It runs itself overnight”about 88 hours wall-clock after the first agents launched; resolution September 5, 2026
“The agents coordinate”2.7 million messages exchanged (Navier–Stokes alone)
“Tokens are cheap”approximately 130 billion output tokens
“It verifies its own work”+17 hours of Lean formalization and verification — on GPT-6 Astra, a different model from the agent fleet
“Millions of dollars in computing power”no dollar figure published by OpenAI anywhere. OpenAI’s chief research officer Mark Chen told reporters costs ran “emphatically in the millions of dollars” (AFP); a third-party estimate reported by Business Insider put the 130B output tokens at about $6.5M and the all-in total at $10M–$40M
“It solved the problem”a claimed, disputed result on an unreleased model, scoped by OpenAI to alternatives (C) and (D) of the Millennium Prize formulation — the forced case — and not accepted by anyone

One scale guard before the numbers travel: those totals are for Navier–Stokes only. Across all the problems the system attempted, OpenAI published 4.9 million messages and about 300 billion output tokens. The two pairs are different measurements; never quote 4.9 million messages alongside 130 billion tokens.

Should you fan out to more agents or keep one agent?

Fan out only when your subtasks are independent, mechanically verifiable, and genuinely parallelisable — all three. Keep one agent when the work is sequential, needs one long shared context, or is expensive to get wrong and hard to check. In practice that means most agency work stays single-agent, and the fan-outs that survive the test are the boring, structured ones.

Fan out only when all three conditions hold

Two practical riders. First, budget the fan-out in agent-hours before you approve it — agents multiplied by hours, not a count of agents. Second, remember that the three conditions are what the anchor run had: subtasks that could be checked independently, a mechanical checker (Lean), and enough parallelism to compress months of mathematics into four days. That is why it was the right shape for a fleet — and it is why your client’s “summarise these 400 documents differently each week” task is usually the wrong shape.

Keep one agent when any of these is true

The arithmetic to run before you approve a fan-out

Convert every candidate fan-out into agent-hours, then price it against an implied rate. The rate below is inferred, not published: it takes the third-party all-in estimate reported by Business Insider ($10M–$40M) and divides it by the run’s 880,000 agent-hours.

Project shapeAgents × hoursAgent-hoursImplied all-in cost (inferred — $11.36–$45.45 per agent-hour)
One agent, one afternoon1 × 44$45–$182
Small crew5 × 840$455–$1,818
Squad20 × 20400$4,545–$18,182
The published anchor (for scale)10,000 × 88880,000$10,000,000–$40,000,000

The implied rate is exactly the third-party band divided by the agent-hours: $10M ÷ 880,000 = $11.36 per agent-hour; $40M ÷ 880,000 = $45.45 per agent-hour. The same division by messages gives $3.70 to $14.81 per message ($10M ÷ 2.7M, $40M ÷ 2.7M). Read that as an order-of-magnitude sanity band around a frontier fleet, not as a rate card — the run was a research sprint on an unreleased model, and OpenAI itself did not publish a price.

Do not extrapolate the published per-agent-hour rates linearly to a bigger fleet. Message volume is what makes fleets super-linear: every message makes a receiving agent re-read context, so input-token consumption grows with the number of agents, while output tokens (the only figure published) stay flat per agent-hour. The honest planning statement is narrower than it looks: agent-hours scale with agents × hours; token cost does not.

Verification is the unmodelled half of the bill

Generation is metered and visible — 2.7 million messages, 130 billion output tokens, 88 hours, all published. Verification is not in the meter: the anchor run published a 17-hour Lean pass with no token count and no cost, and its result is still disputed. That asymmetry is the most transferable thing in the whole run, and it is the line item an agency buyer should put in the contract.

What OpenAI’s 10,000-agent Navier–Stokes run published, and what it did not

Split the two halves and the gap becomes obvious.

PublishedNot published anywhere
on the order of 10,000 concurrent agents (final stage)the exact agent count
about 88 hours to resolutionany cost attributed to OpenAI in its own words
2.7 million messages (Navier–Stokes only)input tokens — at any scope, the single biggest hole in any model of this bill
approximately 130 billion output tokens (Navier–Stokes only)the retry / wasted-token share
+17 hours of Lean formalization and verification on GPT-6 Astrathe token count or cost of that 17-hour verification pass
the model was internal and unreleasedcontext length per agent, GPU-hours, energy cost, and any per-problem token split

The dispute is the evidence

The run is the proof of its own point: the generation half has four clean published numbers, and the verification half has a dispute. Concretely, as of this writing:

For an agency, the transferable lesson is not about mathematics. It is that generation is cheap to meter and verification is expensive to settle — and that a fleet produces more output precisely where the settling cost rises. A 130-billion-token bill is visible on day one. The argument about whether the work was right, whose it was, and who accepts it can run longer than the project — and it is billed in human hours, not tokens.

How should an agency ask for fleet work to be quoted?

Ask for three numbers, always: generation priced per agent-hour, the message volume that drives context re-reads, and verification as its own line item. Never accept a price quoted per agent — the agent is a shape, not a unit of consumption, and a per-agent price hides the hours that actually create the bill.

Line item to demandUnit to insist onWhy
Generationper agent-hour, or per 1,000 messagesagents × hours is the quantity that scales with the bill; a per-agent price lets a vendor answer “how many agents?” instead of “how many hours?”
Message volume and the context-re-read assumptiona message count, plus the assumed input tokens per messagemessages make agents re-read context; input tokens are unpublished for the anchor run, so any quote contains a hidden assumption. Make the vendor state it.
Verificationits own line item, unit = per artefact checked, per test suite run, or per reviewer-hour — and fixedthis is the unmodelled half. It is also the half that decides whether the deliverable is worth anything.
Meteringthe four quantities OpenAI published: agent count (as an order of magnitude), wall-clock hours, messages, output tokensthose four are what made the anchor run auditable. If your vendor cannot report them, you cannot check the invoice.
Stop rule and waste allowancea hard cap on agent-hours, plus a stated retry / wasted-token percentagethe retry rate was never published for the anchor run, so it is an assumption in every quote — and it should be the vendor’s assumption to defend.

Sanity checks on any fleet quote you receive

A worked example

A client wants 400 product pages drafted, each needing an independent check against a source sheet. Say the fleet is 8 agents for 20 hours: 160 agent-hours. At the inferred band that is $1,818–$7,273 of generation — and, at the published-derived rate of 3.07 messages per agent-hour, roughly 490 messages, each of which makes a receiving agent re-read context. Verification is separate: 400 artefacts, one schema check each, priced per artefact. The fan-out passes the three-condition gate (independent, verifiable, parallelisable), so it is the right shape — but if the client cannot state the schema, the third condition fails and the honest answer is one agent, a longer deadline, and a much smaller bill.

For the cost side of the same arithmetic — per-agent-hour and per-message rates, a scaling table for 100, 1,000 and 10,000 agents, and the super-linear fan-out trap — see our sibling anchor on what a 10,000-agent run actually costs.

Frequently asked questions

Should I use multiple AI agents or keep one agent?

Fan out only when the subtasks are independent, mechanically verifiable, and genuinely parallelisable — all three at once. Keep one agent when the work is sequential, needs one long shared context, or is cheap to get wrong and expensive to check, which describes most agency deliverables.

What does an AI agent fleet cost?

No vendor has published a fleet price; the usable anchor is the 2026 OpenAI run, whose 10,000 agents × 88 hours equal 880,000 agent-hours, which the third-party $10M–$40M estimate puts at an inferred $11.36–$45.45 per agent-hour. Full model, scaling tables and the extrapolation trap: what a 10,000-agent run actually costs.

What should a fleet quote tell me about context re-reads?

Because messages make agents re-read context — the published run issued 3.07 messages per agent-hour, and every one of them returns context into a model call. Input-token consumption grows with the fleet while output tokens per agent-hour stay flat. Input tokens were not published for the anchor run at all — which is exactly why the third-party estimate spans $10M–$40M rather than landing on a single figure.

How much did the OpenAI Navier-Stokes run cost?

OpenAI published no dollar figure — its post contains no cost language at all. Our sibling cost anchor prices the workload in full: what a 10,000-agent run actually costs.

Was the OpenAI Navier-Stokes result verified?

No. This result is disputed and unreproducible publicly: it ran on an internal, unreleased model, the MathOverflow discussion states it has not been independently verified, Lean machine-checking is not peer review, and the Clay Mathematics Institute has accepted nothing. Credit is also contested by NYU’s Tristan Buckmaster and Anthropic’s Levent Alpöge.

Sources

  1. OpenAI, “Finite Time Blowup for Navier–Stokes” (post announcing the agents, the 88 hours, 2.7M messages and ~130B output tokens) — https://openai.com/index/navier-stokes-solution/. Note on the route: the live host returned HTTP 403 to three different user agents from our network, so this post was read from the Internet Archive snapshot timestamped 2026-09-08 17:15:18 UTC; OpenAI’s post contains no cost language.
  2. CNBC, September 9, 2026 (agent count and 88-hour quotes; Clay “has not yet commented”; Buckmaster’s questions) — https://www.cnbc.com/2026/09/09/openai-navier-stokes-math-problem-solved.html
  3. phys.org (AFP), September 9, 2026 (Mark Chen: costs ran “emphatically in the millions of dollars”; Clay president Martin Bridson on an “unhurried” process) — https://phys.org/news/2026-09-openai-ai-agents-math-hardest.html
  4. TechSpot, September 2026 (the unreleased model; Noam Brown’s “a very expensive process”; Lean checks steps under assumptions) — https://www.techspot.com/news/113785-openai-claims-10000-ai-agents-solved-one-mathematics.html
  5. Interesting Engineering, September 8, 2026 (2.7M messages, ~130B output tokens, the 17-hour Lean pass on GPT-6 Astra, the Euler side quest) — https://interestingengineering.com/ai-robotics/openai-navier-stokes-mystery-solved
  6. Science Times, September 9, 2026 (Clay’s acceptance rules: peer-reviewed publication plus two years of community acceptance; “not officially considered solved”) — https://www.sciencetimes.com/articles/62566/20260909/openai-claims-have-solved-million-dollar-navierstokes-problem-88-hours.htm
  7. Business Insider, September 2026 (the third-party cost estimate: ~$6.5M for output tokens alone at “OpenAI’s average consumer price”, $10M–$40M including the larger input-token volume) — https://www.businessinsider.com/openai-math-problem-solved-tokens-cost-altman-2026-9
  8. MathOverflow question 515056, September 8–9, 2026 (“It hasn’t been independently verified as of this writing”; the (A)/(B) scope gap) — https://mathoverflow.net/questions/515056
  9. Wikipedia, “Navier–Stokes priority controversy” (the credit dispute timeline) — https://en.wikipedia.org/wiki/Navier–Stokes_priority_controversy

Method note: the published figures above are OpenAI’s own claims about the workload; the derived figures (880,000 agent-hours; 147,727 output tokens and 3.07 messages per agent-hour; the per-agent-hour and per-message dollar bands) are arithmetic on those published totals and on a third-party estimate, and are labelled as derived or inferred wherever they appear. Nothing on this page should be read as a statement that the mathematical result is correct, accepted, or independently confirmed.

Building the client quote? Put the whole fan-out on one page first: agent-hours, message count, and the verification line item that nobody else quotes.

Work the 10,000-agent anchor →