Designing a Production RAG Pipeline
Ingestion, shared embeddings, query rewriting and cited streaming answers — the pieces that turn a RAG demo into a product feature.
By Rahul Gupta · RahulX Labs
Ingestion is half the system
PyMuPDF extraction, semantic chunking and FastEmbed (ONNX) embeddings into Qdrant are the foundation. SHA-256 document hashing deduplicates uploads before you pay for embedding again.
Without race-safe ingest, two users uploading the same PDF can both embed it. A pending → processing conditional UPDATE acts as a distributed lock.
Rewrite for retrieval, answer with the original question
Conversational follow-ups (“why is it useful?”) make bad embedding queries. A small LLM rewrites them into standalone questions for retrieval, while the original question is still passed to the main model so intent is not lost.
Grounding for higher-stakes content
The medical report summarizer uses the same principle in a simpler form: retrieve from the uploaded report before summarising, so answers stay tied to source chunks rather than free-form generation.
Production RAG is less about picking a trendy framework and more about retrieval quality, access control and failure modes you can explain.