Back to Home

Introduction to RAG (Retrieval Enhanced Generations): Connecting Large Models with an 'External Brain'

September 18, 2026 at 08:02 AMSource: RunByAI0 comment(s)TechGuide

The big model knows a lot, but it has two shortcomings that cannot be avoided: first, knowledge has a deadline, and it does not know what happens after training; Secondly, it is not good at remembering the private information it has read. RAG (Retrieval Augmented Generation) is a set of engineering methods proposed to solve these two problems.

1、 What exactly is RAG doing

In summary, before the model answers a question, it first retrieves relevant information from an external knowledge base, and then hands over the information along with the question to the model, allowing it to answer based on the data rather than relying solely on memory.

The complete RAG process typically consists of three stages:

1. Retrieve: Convert the user's question into a vector and find several document fragments in the vector database that are semantically closest.

2. Augment: Concatenate these fragments into context and combine them with the original question to form prompt words.

3. Generate: The large model reads the context, generates evidence-based answers, and can annotate which section of data is referenced.

2、 Why is vector retrieval necessary

Traditional keyword search relies on literal matching, where users ask 'how to make the model not talk nonsense' and the document says' alleviate the illusion of a big model '. The literal names do not overlap, and keyword search is likely to yield nothing.

The method of vector retrieval is to first use an embedding model to map the text into high-dimensional vectors, where semantically similar texts are closer in the vector space. This way, even if the wording is different, it can still be retrieved. This is also the reason why it often appears together with "vector databases".

3、 Typical Value of RAG

-Knowledge can be updated: Newly added data only needs to be re stored, without the need to retrain the model.

-Traceability: The answer can be accompanied by the source for easy verification and to reduce the risk of hallucinations.

-Low cost access to private data: Internal documents of the enterprise do not need to enter the training set, avoiding data leakage and compliance issues.

-Compared to fine-tuning, it is more flexible: updating the search database is much more cost-effective when data changes frequently.

4、 Common engineering difficulties

RAG looks simple, but there are many pitfalls when landing:

-Slicing strategy: Cutting the document too thinly will lose context, while cutting it too large will introduce noise, requiring a combination of paragraph structure and semantic boundaries.

-Search quality: Relying solely on a single vector search can easily miss the results of precise keyword matching. In practice, a mixed approach of vector search and keyword search is often used.

-Reordering: After initially recalling dozens of results, using a reordering model (Reranker) to carefully select can significantly improve the relevance of the final context.

-Illusion residue: Even if given data, the model may still ignore the free expression of the data, which needs to be alleviated through prompt word constraints and answer verification.

5、 How to choose between RAG and fine-tuning

The two are not in opposition. The empirical approach is to inject factual and up-to-date knowledge, and prioritize RAG; If it is necessary to change the expression style, output format, or specific task capabilities of the model, then consider fine-tuning. Many production systems combine the two - first fine tune the model to "understand the rules", and then use RAG to make it "evidence-based".

Summary

The core idea of RAG is simple: instead of letting the model memorize everything, let it learn to look up information when needed. It decouples "memory" from model parameters and hands it over to external retrieval systems for management, thus becoming one of the most mainstream solutions for the implementation of large models.

【 Reference source 】 Comprehensive compilation of industry information and mainstream technical documents that have been publicly released.

large modelRAGSearch enhanced generation
Discussion

Comments (0)

No comments yet. Be the first!

Leave a Comment