Retrieval enhanced generation (RAG) has become the mainstream solution for enterprises to integrate large models into private knowledge bases: first retrieve relevant information, and then have the model answer based on the information to alleviate illusions and supplement private knowledge. But many projects spend almost all their energy on selecting models and tuning prompt words, only to find that the answers are always wrong after going online - the problem often lies not in "generation", but in "retrieval".
Treating RAG as a data pipeline that requires continuous optimization, answering the following five questions before going live can save a lot of rework after going live.
Firstly, does the slicing method respect the content structure? Fixed word count hard cutting will truncate a complete concept, resulting in both loss of context and introduction of noise during retrieval; Splitting by semantic boundaries such as chapters, tables, and code blocks usually results in a more stable hit rate. Different document types (policy systems, product manuals, dialogue records) often require different segmentation strategies.
Secondly, is the search method too singular? Pure vector retrieval is not sensitive to precise information such as keywords, numbers, and models. Mixing vector retrieval with keyword retrieval (such as BM25 algorithm) and then fusing the results is a common practice to improve recall rate, especially suitable for proprietary terms that exist in a large number of enterprise documents.
Thirdly, is there any rearrangement process after the recall? The Top-K results may only have one or two items that are truly relevant. First, expand the recall in a lightweight manner, and then use a rearrangement model to refine and feed the first few items to the larger model, which can significantly reduce the problem of "answering irrelevant questions".
Fourthly, how to manage the updates and versions of the knowledge base? Whether the old content is taken offline in a timely manner after document revision, and whether there will be conflicts between multiple versions, directly determine the timeliness and consistency of the answers. Suggest establishing an update process and effective time record for the knowledge base, rather than throwing old and new files together.
Fifth, is there a quantifiable evaluation set? Organize dozens or even hundreds of "question standard answer" evaluation pairs around real business problems, and any changes (such as changing cutting blocks, models, or adjusting parameters) will be scored first before going online. RAG projects without evaluation sets rely on intuition for optimization and luck for regression.
Answering these questions clearly is more cost-effective than repeatedly patching up prompts after going live. RAG is not a plugin that simply requires a vector library, but rather an engineering chain that is intricately linked from slicing, indexing, recalling, rearranging to evaluation.
[Reference source] Comprehensive compilation of industry information publicly released.