Token Reduction

  • Last Updated 11 August, 2026
  • by David Spuler, Ph.D.

What is Token Reduction?

Token reduction is an LLM inference optimization method that aims to speed up LLMs by reducing the number of tokens processed. This is relevant to any inference process, but is particularly applicable to optimizing multi-step reasoning algorithms, such as Chain-of-Thought, because their interim reasoning steps generate long sequences of tokens.

Token reduction methods have a long history of research in single-step inference optimizations. Some examples where input texts have many tokens include RAG chunks and the conversational history context in chatbot sessions. Hence, reducing the total number of tokens will improve speed in several situations:

One of the simpler ways to reduce tokens is to use prompt engineering techniques. There are two main ways:

  • Write a concise prompt for the LLM (fewer input tokens).
  • Politely ask the LLM to "be concise" as part of the prompt (fewer output tokens).

There are various technical ways to do token reduction as part of inference optimization, and some of the many sub-techniques include:

Token Reduction Optimizations: Book Excerpts and Blog Articles

Free online book excerpts with full text chapters online and free PDF downloads, and the Aussie AI blog, including related articles:

  • Token compression is an LLM inference optimization that reduces an input text sequence to a shorter sequence of text or tokens. There are various methods that can be used before sending the tokens to the LLM, or during the LLM inference processes. Efficiency of inference is improved when there are fewer tokens for the LLM to process... more about Token Compression »
  • Token folding is an LLM inference optimization that "folds" or merges two or more tokens together. This is a type of "token reduction" that results in fewer overall tokens for the LLM to process. Folding of tokens can occur in prompt preprocessing or dynamically during inference computations ... more about Token folding »
  • Token dropping is an LLM optimization that reduces token processing by "dropping" some of them. It is similar to "token pruning", but often refers to dropping tokens during the training phase, whereas token pruning is mostly an inference optimization ... more about Token dropping »
  • Token fusion is an LLM inference optimization that fuses two or more input tokens together. This results in fewer tokens for the LLM to process. Static token fusion occurs prior to LLM processing (during prompt preprocessing), whereas dynamic token fusion occurs during inference (when important tokens can be identified) ... more about Token fusion »
  • Tokenization: The tokenizer does not receive as much attention in the research literature as other parts of large language models. This is probably because the tokenization phase itself is not a bottleneck in either inference or training, when compared to the many layers of multiplication operations on weights. However, the choice of the tokenizer algorithm, and the resulting size of the vocabulary, has a direct impact on the speed (latency) of model inference ... more about Tokenization »
  • Token merging is an LLM inference optimization that merges two or more adjacent tokens into a single token. This optimization reduces the number of tokens for a GPU to process, thereby directly improving speed, and reducing compute. Token merging is a type of token reduction, and is closely related to token pruning, which is simply removing tokens from processing. Merging of multiple tokens together can be performed before sending the prompt to the LLM, or dynamically during inference based on analysis of which tokens are important ... more about Token merging »
  • Token pruning is a type of model "length pruning" that aims to address the cost of processing input sequences and the related embeddings. It is closely related to "embeddings pruning", but is orthogonal to depth pruning (e.g. layer pruning) or width pruning (e.g. attention head pruning) ... more about Token pruning »
  • Token selection is the final phase of LLM processing, also called the "decoding algorithm", whereby the output token is chosen. There are numerous variants including greedy decoding, top-k decoding, top-p decoding, and many more. Most models work by outputting one token at a time, but newer models use "multi-token prediction" (MTP) or "multi-token decoding" to output two or more tokens in parallel... more about Token selection »
  • Token skipping is an LLM inference optimization methods that "skips" the processing of some input tokens. Fewer total tokens to process means less GPU compute cost and faster inference. Variants of this technique include token pruning, token dropping, and token merging ... more about Token skipping »
  • Vision attention optimization is a speed optimization for the inference in Vision Transformers (ViTs). This uses techniques similar to image attention optimization, such as pruning or merging patches or sections of the images ... more about Vision attention optimization »
  • Vision token merging is the combining of two or more areas of an image for faster AI processing. Vision models tokenize the images into segments, which must each be processed by the LLM. Merging them together allows the model to process only a single area, with reduced accuracy ... more about Vision token merging »
  • Vision token pruning is the removal and avoidance of processing portions of an image or video stream. The idea generalizes token pruning in text models, which remove words, to avoid processing some segments of visual images. This is an effective speedup for image processing or machine vision because many of the areas of a large image are not important, and their ... more about Vision token pruning »
  • Vision token reduction is an LLM inference optimization technique for vision models that involves pruning or merging tokens. Tokenization of vision data is based on image tokenization, and can be optimized in various ways. Vision token pruning is avoiding the processing of redundant or unimportant segments of images, such as recurring blank background regions. ... more about Vision token reduction »
  • Refusal Module: The refusal module of an LLM inference stack is the intelligent part of the model that "refuses" to answer certain inappropriate queries. This usually refers to the actual trained intelligence of the LLM to recognize which queries it should refuse, and is trained to generate polite responses. There is a lot of overlap between a "refusal module" and a "prompt shield" that is an add-on component outside the LLM. Also possible are components like heuristic modules that recognize the use of inappropriate words and generate a canned response, thereby completely avoiding any inference cost on such denied user queries.... more about Refusal Module »
  • Prompt compression is the shortening of LLM prompts automatically so that they contain fewer tokens and can be processed more efficiently. This isn't used on short user queries, but there are many other parts of a prompt that are longer, such as the entire conversational history, retrieved documents, or other types of "context" for the prompt. For this reason, this technique is often called "context compression" or "token reduction".... more about Prompt compression »
  • Length pruning is weight pruning on one of the three axes of pruning. The other two axes are width pruning (e.g. attention head pruning) and depth pruning (e.g. layer pruning and early exit). All three types of pruning are mostly orthogonal to each other and can be combined into triple pruning. The main types of length pruning along the "lengthwise" dimension of the inputs are... more about Length pruning »
  • Context compression is an LLM inference optimization that reduces processing of tokens in the context of a query. It is a type of "prompt compression" that involves aspects of techniques such as token pruning or token merging ... more about Context compression »
  • Input compression is an LLM inference optimization that reduces an input text sequence to a shorter sequence of text or tokens for the LLM to process faster. LLM inference efficiency is improved if there are fewer tokens for the LLM to process. ... more about Input compression »
  • David Spuler, Ph.D., Feb 6th, 2026 (updated), 500+ LLM Inference Optimization Techniques, Aussie AI Blog, https://www.aussieai.com/blog/llm-inference-optimization
  • David Spuler, Michael Sharpe, June 2025, Cheaper RAG, Chapter 4, "RAG Optimization: Accurate and Efficient LLM Applications", https://www.aussieai.com/book/rag-book-4-cheaper-rag
  • David Spuler, Ph.D., 16 April, 2026, Chain-of-Thought Efficiency Optimization, Aussie AI Blog, https://www.aussieai.com/research/cot-optimization
  • David Spuler, Ph.D., Dec 21st, 2024, Multi-Step Reasoning Inference Optimization, Aussie AI Blog, https://www.aussieai.com/blog/reasoning-inference-optimization

Research on Token Reduction

Some of the general papers on token reduction strategies:

Token Efficiency in Reasoning and CoT

Token reduction is one of the main methods to improve the efficiency of reasoning models, especially using Chain-of-Thought (CoT) algorithms. There are various ways to improve token counts in CoT at the high level by skipping steps or pruning paths, and there are also various low-level methods.

Blog articles on reasoning efficiency:

More research information on general efficiency optimization techniques for reasoning models:

Efficiency optimizations to Chain-of-Thought that aim to reduce token processing include:

More AI Research

Read more about: