AI & IntelligenceAugust 27, 20266 minutes

Six Practical Ways Employees Can Reduce AI Token Spend

Anna Katharina Bollé Author Profile Headshot
Written byAnna Katharina Bollé
AI & IntelligenceAugust 27, 20266 minutes
Corporate Purchasing Card (P-Card) Guide 

Key takeaways

  1. Tokens are small units of input and output processed by an AI model.
  2. Conversation history, documents, response length, retries, and model choice can all increase spend.
  3. Employees can reduce waste without writing code or avoiding useful AI work.
  4. The aim is a useful result with less unnecessary processing, not the lowest possible token count.

Anyone who opens an AI tool, sends a prompt, or uploads a document starts the token meter. While finance receives the bill later, everyday employee choices can dramatically reduce token usage and reduce AI costs.

A quick explanation of tokens

If you wonder what is an AI token, it is simply a small chunk of text processed by an AI model. It might be a word, part of a word, punctuation, or a number. As a rough English-language guide, 1,000 tokens amount to about 750 words.

The model counts what you send as input tokens and what it returns as output tokens. Input can include your prompt, uploaded material, system instructions, and earlier messages in the conversation. Output is the answer. Providers commonly charge more for output, so both the amount of context you provide and the length of the response matter.

A single prompt may cost very little, but hundreds of employees and automated tools can create usage independently. People therefore need to know which habits add tokens and which ones remove waste.

Here are six practical ways employees can achieve token optimization and reduce token usage.

1. Start a new conversation for a new task

Long conversations can carry earlier messages into every new request. If you finish one task and move to an unrelated one, start a fresh chat instead of keeping the old thread alive.

For example, once you have finished reviewing spend for one cost centre, open a new conversation before drafting a supplier update. The model no longer needs the earlier figures, corrections, and discussion. The new chat uses less input and is easier for the model to follow.

Keep the existing thread when the history is genuinely needed. The point is to stop unrelated context travelling with you out of habit.

2. Include only relevant context and remove repetition

More context does not automatically produce a better result. Share only the information needed for the task, remove duplicated instructions and keep one current version of any material.

For example, if you want help rewriting one section of an email, provide that section and the relevant context instead of the entire email chain and several earlier drafts. Use the full source only when the model genuinely needs to review or compare everything.

3. Set the output length before the model starts

Do not ask for a full report when a short answer will do. State the format and limit in the first request.

You could ask for five bullets, a summary of no more than 150 words, or a short email with two paragraphs. For a data review, request only the exceptions and the reason each one was flagged. These instructions reduce unnecessary output while making the result easier to use.

Avoid asking for a very short response when the task genuinely needs detail. An answer that is too thin may create more correction rounds and use more tokens overall.

4. Make the first request clear and structured

An unclear prompt often leads to clarification, correction, and repeated regeneration. State the task, what the model should focus on, and the expected output format from the start.

For example:

Task: Write a LinkedIn post announcing our new product feature. Focus: The customer problem, the main benefit, and how to try it. Output: Maximum 150 words with a short headline and call to action.

These labels simply separate the different parts of the request. A clear first prompt makes it easier to get a usable answer without several extra turns.

5. Reuse prompts and skills to save time and avoid rework

If you repeat the same task regularly, save the instructions as a prompt template or reusable skill. For example, use one standard prompt to summarise weekly sales calls, identify customer objections, and suggest follow-up actions.

This mainly saves time and keeps outputs consistent rather than directly reducing token usage. However, a tested template can prevent unnecessary explanations, corrections, and regenerations. In API-based workflows, stable recurring instructions may also benefit from lower-priced cached input tokens. Review templates regularly and remove outdated or duplicated instructions.

6. Use the model and speed the task requires

If your company tool allows model selection, using different models like a lightweight approved model works best for straightforward work such as classification, short summaries, or formatting. Reserve premium models for complex or high-stakes tasks that need the extra capability.

The same principle applies to fast or priority processing. Normal processing is enough when the answer is not urgent. Pay for greater speed only when waiting would create a real business problem.

Cheaper is not automatically better. If a weaker model needs repeated corrections, it may cost more per successful task. Judge the result on quality and total effort, not the rate shown beside the model name.

Give employees usable defaults

These habits make token optimization easier to follow when the company provides approved models for common tasks, reusable templates, and simple guidance on response length. This structured approach helps reduce AI costs without inspecting employee prompts or responses.

Employees do not need to memorise provider pricing. They need to understand that every extra page, reply, regeneration, and premium option can add to the bill. A few deliberate choices at the start of a task can prevent a long and expensive conversation later.

FAQs

Anna Katharina Bollé Author Profile Headshot

The Author:

Anna Katharina Bollé

Anna made the shift from working in finance to working on an AI-first product team at Moss. Together with her team, she's now exploring and experimenting with how AI and new ways of working can help finance professionals in their everyday work.