Field Essay
The frontier gap is a definition, not a finding
OpenAI ranked enterprise customers by tokens per active user, called the top ten percent frontier firms, then reported that frontier firms use more tokens. The skills and plugins numbers in the same report are the ones that hold, because their denominator is fixed.
· 4 min read · Kumaresh Bhuyan

OpenAI ranked its enterprise customers by tokens used per person, called the top ten percent frontier firms, then reported that frontier firms use more tokens. That is a definition, not a finding.
The headline number, published on 12 August, is that frontier firms now generate 8.3 times as many output tokens per active user as typical firms, up from 2.6 times in January. The methodology sits a few paragraphs below. Each month OpenAI ranks enterprise customers by output tokens per active user, takes the top ten percent as frontier and the forty-fifth to fifty-fifth percentiles as typical. The gap is the ranking rule restated as a result. It could not have come out any other way.
OpenAI is straight about this, which is worth saying plainly. The report calls tokens an imperfect measure of business value and states elsewhere that there is no single AI adoption leaderboard. The caveat is right there in the text. What happens next is that it falls away in the coverage, and 8.3 times arrives in somebody's board pack as a target.
We already know what that does, because it has been run as an experiment. Uber pushed adoption of AI coding tools through an internal leaderboard ranking teams by total usage, and burned its entire 2026 budget for those tools in four months. In May its chief operating officer, Andrew Macdonald, told Fortune that the link between the spending and shipped customer value was not there yet. By June the company had capped spending at 1,500 dollars per person per month on each tool. In fairness, part of that overrun was a budget set in 2025, before agents consuming tokens at this rate existed. The leaderboard did not cause all of it. It did make sure that nobody slowed down.

The same report carries numbers that are not circular, and those deserve attention. Among weekly active users, 21 percent at frontier firms use plugins against 9 percent at typical firms, and 19 percent use skills against 3 percent. Both are measured as a share of active users on each side, so the denominator is fixed and the comparison holds. A skill is an instruction your organisation has written down once so an agent can reuse it, and a plugin wires that instruction to your real systems. The six-fold difference on skills is not a spending gap. It is a gap in whether a company has encoded how it works in a form a machine can act on, which is a management artefact rather than a purchase.

The strongest objection is that the token gap is not empty. Firms delegating long multi-step work genuinely do generate more tokens, so volume carries real information about depth of use, which is what OpenAI argues. I think that objection is correct. What survives it is narrower and more useful. As a description of deep adoption seen from outside, the measure is informative. As an objective handed to a team, it inverts, because the cheapest way to move it is to send more work to the model whether or not that work needed doing. The report hints at the confound when it puts the gap at 11.7 times in information and technology against 5.3 times in manufacturing.
Here is what I would do with it. I would take the skills and plugins numbers seriously and refuse to let any percentile from it enter an objective. If a team needs an adoption target this quarter, I would set it on the share of people who have a reusable skill wired to a real system, because that number cannot be moved by spending more and it leaves an asset behind when the pilot ends. I would also check whether anything in our own reporting already ranks people or teams by usage, since that is the Uber mechanism arriving quietly and it usually starts life as an enablement dashboard. Where a vendor offers to benchmark us against its own customer base, and OpenAI now offers exactly that, I would take the diagnostic and decline the scoreboard.
None of this requires distrusting the vendor. It requires noticing that a number can be accurate, published with its caveat, and still be the wrong thing to manage toward, because whoever decided what to count has already decided what you will optimise.
Which number in your current AI reporting would improve if your team simply used more of the tool, and what would you have to measure alongside it to tell that apart from progress?
The Execution Edge
A monthly deep dive on enterprise AI and turning strategy into execution, written from inside the programme office rather than above it.