A RAG system is two pipelines. One prepares your documents so they can be searched. The other answers questions by searching those documents and passing the best passages to an LLM. Most quality problems start in the first pipeline, even though people notice them in the second.
PipelinesPrepare, then answer
SearchVector plus keyword
AnswersWith sources
PermissionsChecked at search time
Reference architecture
The two pipelines
Scroll sideways to see the whole diagram →Left: documents are prepared continuously as they change. Right: each question is answered from the prepared index.
Part
What it does
1 Collect
Connectors pull documents from where they live, with their permissions and last-modified dates.
2 Parse
Text, tables and scanned pages are converted cleanly. This step decides more about quality than the choice of LLM.
3 Split and embed
Documents are split into passages that keep their headings, and each passage gets an embedding for meaning-based search.
4 Search
The question is matched by meaning and by keywords, and only passages the user may read are returned.
5 Rerank
A second model reorders results by relevance, and current documents are preferred over archived ones.
6 Answer
The LLM answers only from the passages it was given, cites them, and says when the answer is not there.
Typical tools
Layer
Common choices
Parsing
Unstructured, Docling, Apache Tika, OCR for scans
Embeddings
Open multilingual embedding models (for example BGE or E5 families) or hosted embedding APIs
Search index
OpenSearch or Elasticsearch, PostgreSQL with pgvector, Qdrant, Milvus
Reranking
Cross-encoder rerankers
Orchestration
LlamaIndex, LangChain or a small amount of custom code
Tell us where the documents live, who should be able to ask, and ten questions people ask today. We will come back with a plain view of what it takes to answer them well, and safely.