AI Engineering8 min

Designing a Production RAG Pipeline

Ingestion, shared embeddings, query rewriting and cited streaming answers — the pieces that turn a RAG demo into a product feature.

By Rahul Gupta · RahulX Labs

Ingestion is half the system

PyMuPDF extraction, semantic chunking and FastEmbed (ONNX) embeddings into Qdrant are the foundation. SHA-256 document hashing deduplicates uploads before you pay for embedding again.

Without race-safe ingest, two users uploading the same PDF can both embed it. A pending → processing conditional UPDATE acts as a distributed lock.

Rewrite for retrieval, answer with the original question

Conversational follow-ups (“why is it useful?”) make bad embedding queries. A small LLM rewrites them into standalone questions for retrieval, while the original question is still passed to the main model so intent is not lost.

Grounding for higher-stakes content

The medical report summarizer uses the same principle in a simpler form: retrieve from the uploaded report before summarising, so answers stay tied to source chunks rather than free-form generation.

Production RAG is less about picking a trendy framework and more about retrieval quality, access control and failure modes you can explain.