What is a knowledge graph
Knowledge Graph uses a triplet of "entity relationship entity" to organize scattered knowledge into a searchable and inferential network. For example, (Beijing, the capital, China) is the simplest triplet: nodes are entities, edges are relationships. Unlike traditional databases that lock data in tables, graph structures naturally support "multi hop" queries - starting from A, passing through B, and finding C.
In engineering, knowledge graphs are often represented in the form of RDF (Resource Description Framework) or Property Graph, and are accompanied by query languages such as W3C's SPARQL and Cypher commonly used for property graphs.
Why does the era of big models still require knowledge graphs
The big language model is very good at "speaking", but often cannot remember the facts and even fabricates them seriously. The knowledge graph precisely fills three gaps:
Structured: Facts are explicitly expressed, traceable, and interpretable.
• Inference: Derive implicit relationships from known triplets using relationship rules or graph algorithms.
Anti illusion: Provide reliable and traceable external knowledge sources for generative models.
As early as 2012, Google released its own knowledge graph, using the "knowledge panel" to directly answer factual questions such as "what is so and so", which was also the starting point for this term to enter the public eye.
How is a knowledge graph built
Building a knowledge graph is usually a pipeline:
Ontology design: First define entity types and relationship types, and build a "skeleton".
Knowledge extraction: Extracting entities and relationships from unstructured text, commonly using Named Entity Recognition (NER) and Relationship Extraction (RE).
Knowledge Fusion: Align the same entity from different sources, perform disambiguation and deduplication.
Knowledge completion: Using links to predict missing relationships, graph embedding methods such as TransE and Rotate are commonly used tools.
Quality governance: Annotate sources, resolve conflicts, and continuously update.
Graph Neural Networks and Knowledge Graph
Graph neural networks (GNNs) learn node representations by passing messages on a graph, making them naturally suitable for tasks such as node classification and link prediction; Knowledge graph embedding maps entities and relationships to a low dimensional vector space, where 'distance' represents semantic relevance. The two often complement each other: graphs provide structure, while learning models provide reasoning and generalization abilities.
The combination of knowledge graph and big model: GraphRAG
The hottest combination point at present is Retrieval Enhanced Generative (RAG). Traditional RAG relies on vector retrieval, which is often inadequate for problems that require "global understanding" and "multi hop inference"; By introducing knowledge graphs into retrieval, relevant subgraphs can be retrieved based on relationships and then handed over to the large model to generate answers. The recently proposed ideas such as GraphRAG combine the structural advantages of graphs with the expressive power of generative models for global question answering and summarization.
Conversely, large models can also help build knowledge graphs - automatically extracting, completing, and aligning, significantly reducing the cost of graph construction and forming a cycle of "human-machine co construction".
one-sentence summary
The value of a knowledge graph lies not in its size, but in its ability to clarify relationships. When a large model requires trustworthy, traceable, and interpretable knowledge, a structured knowledge network is often more effective than having more parameters.
[Reference source] Comprehensive compilation of industry information that has been publicly released, including W3C open standards such as RDF/SPARQL, public introductions to Google Knowledge Graph, and publicly released technical materials such as GraphRAG.