Scenario
In our work with regional leaders, we have found that many measure AI usage by counting tokens used by employees and departments. Leadership often assumes that higher token counts reflect greater AI adoption.
At first glance, this may seem logical, since higher token usage often produces outputs that wouldn't have been produced otherwise, such as transforming spreadsheet data into an interactive, locally hosted webpage.
This article will help you better understand tokens and how they work in the AI realm, while also offering approaches to increase AI adoption and deliver significantly improved results.
Understanding Tokenisation
In large language models (LLMs), a token is a unit of text the model processes and understands. Tokens are typically made up of a few characters, such as a word, a part of a word, or even punctuation marks. For example, the word "cat" is one token, while the phrase "unbelievable" might be split into several tokens like "unbe", "liev", and "able".
For demonstration purposes, the examples will show tokens being sent to and from an LLM. We’ve tokenised at word boundaries and simplified the returned data for clarity. However, in real-world use, tokens aren't constructed or communicated exactly this way.
Nearly all interactions with an LLM happen through tokens, and providers typically distinguish between input tokens (text sent to the model) and output tokens (text generated by the model). Output tokens generally incur a higher cost than input tokens, as pricing for output tokens often includes both the computational resources required to generate the response and the value of the information produced. As such, understanding this distinction is essential to accurately track usage and manage costs in LLM applications.
Understanding Reasoning
In earlier LLMs, the output was closely tied to a single answer. The model would usually make a one-shot attempt to provide a response. For simple questions, this approach often worked well, since the statistically more likely answer is usually the most acceptable answer. However, as problems grow more complex, the chance of a correct answer on a single attempt decreases.
A more effective approach is to give the model room to work before it answers. Rather than committing immediately, the model produces a long sequence of intermediate steps: breaking the problem down, testing an approach, noticing when a step fails, and trying another. Only then does it commit to an answer. This is what the industry calls reasoning. The work is usually hidden from the user, but it shows up as output tokens and gets billed as such. A reasoning model answering a difficult question can spend the overwhelming majority of its output on thinking that nobody ever reads.
The animation illustrates that reasoning tokens are normally an order of magnitude larger than response tokens. Regardless of the model you choose, most of the response will be reasoning tokens. This is only half the story, as we will see in a moment.
The cost of reasoning
Reasoning presents a challenge. While it can improve the likelihood of a better answer, it is not always clear how much reasoning an LLM should perform to achieve the best result. A higher token count might suggest more effort, but this does not always translate to better quality.
Consider two people who ask the same question to two identical LLMs. It seems reasonable to assume that the person who allows the LLM more reasoning time, or more effort, will receive a better answer. In reality, this is not always the case.
LLM providers understand this tendency. They often design interfaces that nudge users toward higher-effort options that use more computational resources, even for simple tasks. Progress bars and similar visual cues can encourage users to select more complex processing, which boosts token usage.
However, extra reasoning frequently produces the same answer, since most problems have a limited space for reasoning. After a certain point, additional effort does not improve the result and simply wastes resources. LLM providers rarely inform users about the optimal level of reasoning for their task. This leaves users guessing how much is enough.
Understanding context and how it increases token usage
LLMs are stateless. Stateless means they don't remember your previous requests. They process thousands of requests per second. If an LLM tried to remember every request, it would quickly become overwhelmed. Being stateless is an advantage, as every request is treated as if it were the first. The responsibility for remembering the conversation falls to the caller, typically your computer.
This means the caller must remember the entire conversation and transmit it each time a new prompt is sent.
This detail is easy to overlook because the experience hides it. When someone types a follow-up question, it looks like they're sending only one line. Behind the scenes, your computer first sends all your previous questions, then every previous answer, and finally the new prompt at the end.
This effect builds up. The longer the conversation, the larger the input tokens get per request. This means that over an entire conversation, the cumulative total of output tokens grows steadily with each exchange. However, the input side behaves differently. Each turn repeats everything that came before, so the opening question is sent on the first turn, then again on the second, the third, and so on, for as long as the session continues. In long conversations, early exchanges are resent many times. This effect is even greater for models that support media like images, audio, or video.
For anyone reading a token report, this creates a challenge. A long working session will cost much more than a short one covering the same material, but the difference is not obvious to the user. Two colleagues who ask the same six questions, one in a single conversation and the other across six separate sessions, will generate very different token counts for the same work. Neither worked harder than the other.
How reasoning has shifted the token balance
The previous sections illustrate a significant shift that has occurred over the past two years.
Model capabilities have improved iteratively. Modern LLMs are now trained to use a wider range of tools available on a user’s computer, allowing them to read, edit, sort, and locate information using existing software. In addition, each subsequent generation benefits from training on the questions and answers processed by earlier models, a process known as distillation. This has made LLMs more adept at identifying likely correct solutions.
Consequently, recent models require significantly less user intervention. They are capable of independently planning, decomposing complex tasks, and executing them within a single interaction. While fewer interaction rounds might suggest lower token usage, these advanced models allocate more processing power to reasoning before generating an answer. This computational effort is not limited to output tokens. It also shows up in input tokens. Even when users are no longer prompting each step, the entire context of autonomous decisions made on the local machine must still be transmitted to the LLM. As a result, most communication between user and model consists of reasoning content.
As a result, the balance has flipped. What was once mostly input is now mostly output. The same task can create similar token counts in two very different ways.
For context, this reversal is typical for conversational work. Agentic use, by contrast, tends to be input-heavy, since every file, search result, or tool response is sent back into the model as input. For example, a coding agent working through a repository will have input-heavy sessions. The input-to-output ratio depends more on the style of work than on the model itself. This is a key reason to be cautious about interpreting token numbers at face value.
The effect of capabilities on cost
It's reasonable to assume that as models became more capable, they also became more expensive, and that was broadly true for a while. It stopped being true, and understanding why matters for anyone budgeting this work.
The chart illustrates the cost of one million tokens at the time each model was launched, plotted against that model’s performance on a fixed science benchmark that has remained unchanged for four years. Each line shows a single provider's trajectory over time. To ensure a fair comparison, we include only the most cost-effective models capable of meaningful work, rather than flagship releases. This mirrors the decision-making process of organisations focused on managing expenses. Further details on these trends are provided in Appendix A.
From late 2024 onward, the relationship between price and capability diverges. The cost per million tokens continues to decline, while benchmark scores rise sharply. For every major provider, the most affordable models no longer require significant compromise. This shift stems from two distinct developments. First, advances in model design and operational efficiency have lowered the cost of producing tokens. Second, improvements in reasoning have enabled models to solve problems more effectively by working through intermediate steps before delivering an answer. These developments occurred independently.
The second trend introduces a cost that is not reflected in this chart. Because the horizontal axis measures only token price, it does not account for changes in token volume.
Token consumption increased dramatically. For the same model and set of questions, enabling reasoning used about sixteen times more tokens than reasoning-disabled scenarios. More broadly, reasoning models have been observed to require up to twenty times as many tokens as one-shot models. Even models without explicit reasoning have become more verbose across generations.
Taken together, the reality is less reassuring than either chart suggests on its own. While the price per token has dropped significantly, token consumption has risen sharply. Together, this means actual costs have not declined nearly as much as a falling token price alone would suggest.
When tokens matter, and when they do not
This does not make token counts irrelevant. Instead, they answer a very specific question and are often misapplied to others.
For cost-related matters, tokens provide the most accurate measurement. The bill is determined by multiplying token volume by price, making tokens indispensable for forecasting expenses and planning capacity within rate limits. Token counts also work well for comparing models, as long as the comparison focuses on the cost to complete a task rather than headline rates, which can differ significantly. Token usage data is also valuable for identifying inefficiencies. Teams using advanced models for routine work or consistently operating with maximum reasoning effort will be clearly visible in this data. Likewise, teams that have not yet adopted AI can be easily identified. An important consideration when allocating enablement resources.
However, token counts become limiting when used as proxies for adoption, productivity, or value. The metric fluctuates with model architecture, configuration, and the number of conversational turns, none of which reflect genuine user engagement or progress. Teams that solve problems efficiently may generate fewer tokens than those struggling with the same issues, potentially inverting the intended meaning. Importantly, token counts are not linked to tangible outcomes.
Another challenge is that no refinement of this metric can address the underlying issue. Token counts cannot be incorrect, nor can cost figures. They record activity. Activity that did in fact occur. This reliability makes them appealing for reporting but also explains why they are unsuitable for decision-making.
How do you measure true AI adoption & tangible benefits?
To determine whether AI initiatives are truly effective, the assessment must be based on outcomes that could have occurred differently.
A more robust approach is to shift from measuring engagement to systematically collecting and evaluating hypotheses. Any individual, team, or temporary working group can propose a hypothesis. Each hypothesis should clearly specify an owner, the company objective it supports, the baseline measurement before any intervention, the expected change, the metric that will indicate success, and a target date for evaluation. If a hypothesis doesn't align with a defined company objective, don't open it. Keep the process simple to encourage participation. If it requires more than these six elements, it is unlikely to be adopted.
Among these elements, the baseline is critical. Without a pre-intervention measurement, the process risks devolving into retrospective justification, where improvement is claimed without credible evidence of impact. Recording the baseline before any changes provides a practical substitute for a formal control group, and it is what distinguishes rigorous evaluation from internal promotion.
It is equally important to foster a culture where reporting unsuccessful hypotheses is routine. Teams often highlight only their successes, which obscures the true success rate and skews the organisational learning portfolio. Unsuccessful hypotheses should be celebrated almost as much as successful ones. If there is any penalty for reporting failure, teams will avoid proposing ambitious initiatives, and valuable lessons will be lost.
A common leadership concern is the need for timely metrics rather than waiting for long-term outcomes. This is a valid point. Telemetry and activity data are available immediately, whereas outcome-based evidence takes longer to accrue. This time lag is one reason organisations often default to token counts. However, withholding interim metrics is not a solution, as this only encourages others to revert to less meaningful measures. A more effective interim approach is to report on the portfolio of active hypotheses: the number open against each company goal, average time to verdict, invalidation rates, and confirmed value delivered in business-relevant terms.
The invalidation rate warrants particular attention. If it is consistently near zero, it suggests teams are proposing only initiatives they are confident will succeed, meaning the process measures risk aversion rather than genuine capability.
Close
Counting tokens is entirely appropriate. Token data reveals AI costs, identifies inefficiencies, and highlights who has not yet engaged. Each of which provides valuable insight.
The challenge emerges when token counts are presented as proof of successful investment. An increase in token usage could indicate more complex problems, a more advanced model, longer interactions, or simply a configuration change. These scenarios are fundamentally different, but a dashboard will display them in the same way.
Ultimately, decision-makers care about whether these initiatives create meaningful impact. Only outcome-based measures can answer that, and only if a clear baseline was established at the outset.