What is Retrieval-Augmented Generation (RAG)?
An architectural pattern that supplements an LLM query with factual excerpts retrieved from private or external vector databases, preventing hallucinations and ensuring up-to-date answers.
Key facts
- Separates storage of dynamic factual knowledge from the frozen model weights.
- Significantly cheaper and faster to maintain than continuously fine-tuning models.
- Provides verifiable citations and provenance directly to source documents.
- Advanced variations include GraphRAG, Hybrid Search (Dense + BM25), and Agentic RAG.
Explanation
Standard LLMs possess knowledge bounded by their training cutoff and can hallucinate on proprietary corporate data. RAG overcomes this by converting enterprise documents into dense mathematical embeddings, storing them in high-dimensional vector databases, and fetching the top semantic matches whenever a query is issued.
Modern enterprise RAG systems integrate cross-encoders for re-ranking, query expansion, and document chunking strategies to minimize latency and guarantee retrieval fidelity.