Home / RAG / How RAG is built
RAG guide

How RAG is built

A RAG system is two pipelines. One prepares your documents so they can be searched. The other answers questions by searching those documents and passing the best passages to an LLM. Most quality problems start in the first pipeline, even though people notice them in the second.

PipelinesPrepare, then answer
SearchVector plus keyword
AnswersWith sources
PermissionsChecked at search time
Reference architecture

The two pipelines

Scroll sideways to see the whole diagram →
Preparing documents (runs continuously)Answering a questionDocument sourcesSharePoint, file shares, wikis, ticketsParse and cleanPDF, Word, tables, scans (OCR)Split and embedpassages with title, date, permissionsSearch indexvector and keyword, with access listsQuestionplus who is askingRetrievehybrid search, filtered by permissionRerankbest and most current passages firstAnswerLLM writes from the passages, with sourcesUsersask in plain language123456
Left: documents are prepared continuously as they change. Right: each question is answered from the prepared index.
PartWhat it does
1 CollectConnectors pull documents from where they live, with their permissions and last-modified dates.
2 ParseText, tables and scanned pages are converted cleanly. This step decides more about quality than the choice of LLM.
3 Split and embedDocuments are split into passages that keep their headings, and each passage gets an embedding for meaning-based search.
4 SearchThe question is matched by meaning and by keywords, and only passages the user may read are returned.
5 RerankA second model reorders results by relevance, and current documents are preferred over archived ones.
6 AnswerThe LLM answers only from the passages it was given, cites them, and says when the answer is not there.

Typical tools

LayerCommon choices
ParsingUnstructured, Docling, Apache Tika, OCR for scans
EmbeddingsOpen multilingual embedding models (for example BGE or E5 families) or hosted embedding APIs
Search indexOpenSearch or Elasticsearch, PostgreSQL with pgvector, Qdrant, Milvus
RerankingCross-encoder rerankers
OrchestrationLlamaIndex, LangChain or a small amount of custom code
ModelSee enterprise LLM

If you already run OpenSearch, Elasticsearch or PostgreSQL, you probably do not need a separate vector database to start.

Want answers from your own documents?

Tell us where the documents live, who should be able to ask, and ten questions people ask today. We will come back with a plain view of what it takes to answer them well, and safely.