1、 What is word embedding
The first step in making the big model understand text is to turn it into numbers. But using numbers directly (such as "cat=1, dog=2") is not feasible - there is no semantic relationship between the numbers, and the model cannot see that both "cat" and "dog" are animals.
Embedding solves the problem of mapping each word, sentence, or even the entire article into a fixed length vector of real numbers, such as [0.12, -0.83, 0.47,...]. This string of numbers is its' coordinates' in high-dimensional space.
2、 Why can vectors express meaning
The key lies in the training objective: to bring semantically similar words closer together in the vector space. A commonly cited classic example is:
Vector (King) - Vector (Man)+Vector (Woman) ≈ Vector (Queen)
This indicates that the vector not only records the word itself, but also captures abstract relationships such as "gender" and "monarchy". To measure the similarity between two vectors, cosine similarity is usually used - the smaller the angle, the closer the meaning.
3、 From word to sentence: Embedding hierarchy
-Word level: Word2Vec and GloVe were early representatives that learned a relatively fixed vector for each word.
-Context level: Transformers such as BERT and GPT dynamically generate vectors based on context, and the vectors for "apple" in "eating apple" and "Apple company" are not the same.
-Sentence/paragraph level: Compressing the entire sentence into a vector, commonly used for retrieval, clustering, and deduplication, is the foundation of RAG system retrieval.
4、 What does it do in the actual system
-Semantic search: Convert both queries and documents into vectors, and find the most relevant content based on similarity, rather than blindly picking keywords.
-RAG retrieval: vectorize the knowledge base into chunks, and recall the most relevant fragments to feed to the large model when users ask questions.
-Recommendation and clustering: Transform users, products, and articles into vectors and measure their similarity using distance.
-De duplication and classification: Classify similar vectors into one category or determine whether two pieces of text are duplicated.
5、 A common misconception
Many people think that the higher the vector dimension, the better. In fact, higher dimensions can hold more information, but they also consume more storage and computing power, and the returns will decrease beyond a certain point. When choosing an embedded model, it is more important to consider its actual retrieval performance on your task (Chinese, code, or long text?), rather than just focusing on dimensional numbers.
6、 Summary
Word embedding is a crucial step in translating "semantics" into "geometry": text becomes vectors, meaning becomes distance. By understanding this, one can understand why large models can perform semantic search, RAG retrieval, and recommendation - they are essentially "finding neighbors" in a high-dimensional space.
【 Reference Source 】 Comprehensive compilation of industry information released publicly