When we input a sentence into a large model, what the model actually sees is not text, but a string of numbers. The component that converts between "text seen by humans" and "numbers calculated by models" is the tokenizer. Although it is rarely mentioned, it profoundly affects the capability boundaries, costs, and many seemingly "inexplicable" errors of large models.
1、 Why do we need word segmentation
Neural networks can only handle numerical values, so text must first be cut into individual units and then mapped into vectors. If we use "whole words" as the unit, the vocabulary will expand infinitely, and rare and new words will not be represented; If measured in units of 'single characters', the sequence will become very long and the semantics will be too sparse. The compromise approach is to divide subwords: common words are retained as a whole, and rare words are broken down into smaller fragments.
2、 Mainstream word segmentation algorithm
BPE (Byte Pair Encoding): Starting from characters, repeatedly merging adjacent symbol pairs with the highest frequency of occurrence is a commonly used scheme in the GPT series.
WordPiece: BERT adopts a scheme that selects the segments to be merged based on likelihood, rather than simply looking at frequency.
SentencePiece: treats spaces as regular characters without relying on pre segmentation, making it more friendly to languages such as Chinese and Japanese that do not have space separation. It is used in models such as T5 and Llama.
Byte level solution: Convert text into bytes before processing, ensuring that any character can be represented and avoiding the issue of "out of bounds" (OOV) as much as possible.
3、 Why is Chinese more 'fee token'
In English, a large number of commonly used words become a token as a whole, while in Chinese, one Chinese character often occupies one or even multiple tokens. The result is that for the same content, the number of tokens consumed in Chinese is usually significantly higher than in English. This will directly affect two things: how much content can be loaded in the context window, and the cost of calling based on token billing.
4、 The 'strange phenomenon' caused by word segmentation
Models are not good at counting how many letters there are in a word because letters are often mixed together at the token level.
Simple arithmetic is prone to errors because numbers may be cut into inconsistent segments, such as numbers with different digits corresponding to completely different cuts.
Multiple spellings, capitalization, and space differences of the same word may correspond to completely different token sequences.
5、 Inspiration for Users
When writing prompt words, clear and structured expression is usually more effective than deliberately "carved" wording.
When dealing with long Chinese texts, pay attention to the token budget and streamline redundant content if necessary.
When the model makes seemingly "inexplicable" errors, you can first think about whether the word segmentation has shredded the information.
Conclusion
The tokenizer is the most advanced and easily overlooked part of the large model pipeline. Understanding it can not only explain many "strange phenomena", but also help us design prompt words, evaluate costs, and determine the boundaries of model capabilities more reasonably.
[Reference source] Comprehensive compilation of industry information released publicly (including BPE, WordPiece, SentiencePiece, and other publicly available technical documents and open-source implementations).