Build RAG you can trust in production.
RAG is retrieval-augmented generation: an LLM that answers from your documents instead of from its training memory.
- Setting Up
- How to Use This Book
- The Problem RAG Solves
- Build a RAG
- Anatomy of a RAG System
- Your First RAG
- Watching Your Naive RAG Fail
- Embeddings
- What an Embedding Is
- Choosing an Embedding Model
- The Geometry of Meaning
- Documents
- Why Documents Aren't Ready for Retrieval
- Chunking, the Most Underrated Decision
- Metadata, the Free Win Most People Skip
- Storage
- From Lists to Indexes
- Choosing a Vector Store
- Updating an Index Without Breaking Production
- Retrieval
- Pure Vector Retrieval Is Not Enough
- Hybrid Retrieval Done Right
- Query Transformation
- Reranking, the 80/20 Win
- Filtering Without Killing Recall
- Generation
- Stuffing the Context Window
- Prompting for Grounded Generation
- Citations and Attribution
- Streaming Without Losing Citations
- Evaluation
- Why "It Looks Right" Is Not Evaluation
- Building a Golden Dataset
- Retrieval Metrics That Matter
- Generation Metrics: Faithfulness, Relevance, Completeness
- The Eval Loop
- Production
- Latency Budgets
- Cost Optimization
- Observability
- Multi-Tenancy and Access Control
- Failure Modes and Graceful Degradation
- Advanced
- Hierarchical and Multi-Hop Retrieval
- GraphRAG and Knowledge Graphs
- Self-RAG and Corrective RAG
- Agentic RAG
- Multi-Modal RAG
- Capstone
$39 USD
Read the free previewSee pricing