AI & IntelligenceAugust 27, 20266 minutes

AI Token Spend Is Becoming a Finance Problem

Anna Katharina Bollé Author Profile Headshot
Written byAnna Katharina Bollé
AI & IntelligenceAugust 27, 20266 minutes
16 Essential Software Tools for CFOs

Key takeaways

  1. Tokens are small units of text processed or generated by an AI model. Both the request and the response can be billable.
  2. The actual AI token cost moves with model choice, context length, output length, call volume, and automated retries, so they are harder to forecast than software seats.
  3. The invoice may show the provider and amount but still leave finance unable to identify the team, workflow, or business purpose behind the spend.
  4. Effective AI cost optimization needs two layers: finance governance for accountability and engineering and employees control for the way tokens are spend.

The bill goes up. Headcount is flat, and procurement has not approved a new supplier. The cause these days actually may be a model upgrade, longer responses, repeated conversation history, or an agent making dozens of calls for one task.

AI token spend behaves less like a fixed subscription and more like a metered operating cost. Finance needs enough token literacy to explain the bill and connect spend with value.

What is an AI token?

A token is a small unit of data processed by a language model, representing the core metric of LLM pricing. It may be a word, part of a word, punctuation, or a number. As a rough guide, 1,000 tokens amount to about 750 English words. The exact count varies by language and content.

There are two basic categories:

  • Input tokens cover what is sent to the model, including the prompt, uploaded material, system instructions, and relevant conversation history.
  • Output tokens cover what the model generates in response, influencing the overall LLM token cost.

Providers usually price these categories separately, with output commonly carrying the higher rate. The model generates a response step by step, so a long report can cost more than a short classification task even when both start with a similar prompt. Some consumption is less visible. Models may use internal reasoning tokens before producing an answer, and long chats may resend earlier messages on every turn. The user sees one response while the billing system records more processing.

AI adoption is moving faster than cost visibility

AI investments are already substantial and still rising. UK midmarket and enterprise businesses reported average annual AI spending of £15.94 million and expected their investment to increase by a further 40% over the following two years (SAP and Oxford Economics, 2025).

AI use among surveyed UK businesses also rose from 25% in 2024 to 54% in 2026 (British Chambers of Commerce, 2026). Yet only 31% of surveyed businesses already using AI reported a positive return on their investment (Studio Graphene and Censuswide, 2026).

That makes proactive AI cost management increasingly important. Unlike traditional software with a predictable subscription fee, generative AI is often billed by consumption. Every prompt, document upload, automated call, or repeated loop can add to the bill, allowing costs to rise before finance can see or control them.

Why token billing is harder to manage than SaaS

Traditional SaaS gives finance familiar units: seats, plans, renewal dates, and contracted fees. A higher bill usually comes with a new user, an upgraded plan, or a supplier change.

Token billing is driven by behaviour inside the service. The bill can change because:

  • employees send longer prompts or request longer answers
  • a workflow moves to a more expensive model, affecting overall LLM cost management
  • the same conversation history is processed repeatedly
  • an automated agent adds planning steps, tool calls, or retries
  • usage grows after a product release, reporting cycle, or internal rollout

With consumption billing, cost is rate multiplied by usage, and both can change. Providers set different prices for input, output, model tiers, cached input, and processing modes. Teams can also move between models without a procurement event. Finance may only notice when the invoice lands.

AI in a customer-facing product may sit in cost of sales and influence gross margin. Internal assistants may sit in operating expenses. Combining them hides whether the cost supports revenue, reduces internal effort, or is simply drifting.

The invoice rarely explains the variance

AI spend is fragmented across model providers, cloud platforms, subscriptions, coding tools, and automated workflows. The charges appear in different systems and at different levels of detail.

An invoice can identify the supplier and amount. It may not identify the team, cost centre, workflow, agent, model, or API key behind the usage. Without that context, finance cannot answer basic questions:

  • Did the price change, or did usage increase?
  • Which team or workflow caused the movement?
  • Did the business receive more output or merely use a more expensive technical path?
  • Should the spend sit in cost of sales or operating expenses?
  • What should be forecast for the next month?

Useful reporting needs usage metadata alongside the transaction. Finance does not need to read prompts or responses. It needs spend by provider, model, team, workflow, agent, and API key, with a named business owner.

AI cost control works on two layers

Token costs have a technical origin and a financial consequence, requiring collaborative LLM cost optimization. One team cannot manage both on its own.

The engineering & employee usage layer controls model selection, context and output length, caching, routing, batching, rate limits, and retry logic. Engineering can explain why a workflow became more expensive and change its technical path.

The finance and governance layer controls ownership, budgets, cost allocation, showback or chargeback, variance analysis, anomaly alerts, forecasts, and measures of value. Finance decides where a limit belongs and whether the outcome justifies the cost.

That is what makes AI cost control difficult: the people who start the meter and influence how fast it runs are not the ones responsible for keeping spend in check. Any employee can trigger additional costs, usage is highly variable, and finance has no immediate control over either.

Finance needs token literacy before token optimisation

The first task is not to minimise every token. Low usage can be as wasteful as high usage if the work produces no value. The immediate finance task is to make the spend explainable.

Finance should first separate price changes from usage changes, then identify where the movement comes from and whether it comes from valuable growth or inefficiency. Budgets and controls can then follow the activity and patterns that create the expense.

AI token spend can scale quietly and alter margins before traditional software controls catch it. Finance does not need every engineering detail. It needs to know what was consumed, at what rate, for which purpose, and with what result.

FAQs

Anna Katharina Bollé Author Profile Headshot

The Author:

Anna Katharina Bollé

Anna made the shift from working in finance to working on an AI-first product team at Moss. Together with her team, she's now exploring and experimenting with how AI and new ways of working can help finance professionals in their everyday work.