RAG: Retrieval-Augmented Generation

RAG is an application-controlled pipeline that retrieves external evidence, places selected material into current context, and asks a model to generate from it without rewriting the model's parameters.

  • Explainer
  • 6 min read
Illustration of the retrieval-augmented generation pipeline fetching external evidence into model context.

A base model may predate a company’s current policy and have no native knowledge of its private documents. It can still answer a question about that policy when an application finds relevant source material and supplies it with the request.

Retrieval-augmented generation, or RAG, is a family of systems built around that path. The application retrieves external information, adds selected evidence to the model’s current context, and asks the model to generate a response conditioned on it.

RAG does not upload knowledge into the model’s parameters. It changes the evidence available for this inference.

A normalized RAG pipeline

The application controls every handoff around generation
Before requestsPrepare a searchable corpus
  1. External sources / 01Collect

    Documents, records, code, or other maintained source material

  2. Application / 02Chunk

    Create retrievable units while preserving useful metadata and provenance

  3. Retrieval system / 03Represent and index

    Build a searchable structure, often with embeddings, keywords, or both

For one requestRetrieve, supply, generate, check
  1. Request / 04Form query

    Represent the user's information need for search

  2. Retrieval system / 05Rank candidates

    Search, filter, and optionally rerank external material

  3. Application / 06Select and inject

    Place chosen source text inside the current model context

  4. Model / 07Generate

    Produce an answer conditioned on parameters and supplied evidence

  5. Application / 08Check support

    Assess evidence fit, grounding, and citation integrity

What changed for this callExternal evidence entered current context; the base model's parameters did not change
RAG creates an evidence path, not a truth guarantee.Corpus coverage, chunking, ranking, selection, context limits, source quality, model interpretation, and support checks remain distinct control and failure surfaces. Implementations may use dense, sparse, hybrid, or tool-mediated retrieval.

Two phases, one evidence path

A production implementation can vary widely, but a normalized RAG system has two broad phases.

Prepare material for retrieval

Before a user asks a question, the system usually prepares a corpus: the documents, records, tickets, code, or other sources that retrieval is allowed to search.

Preparation can include:

  1. collecting and parsing sources;
  2. splitting them into retrievable chunks;
  3. attaching titles, dates, permissions, tenant IDs, or other metadata;
  4. creating searchable representations;
  5. writing those representations and source references to an index.

Embeddings are common, but RAG is not synonymous with vector search. A retriever can use keywords, dense vectors, structured filters, a search engine, an API, a database query, or a hybrid of several methods. The defining step is that external material is retrieved and supplied for generation.

Retrieve for one request

At request time, the system turns the user’s information need into a query, searches the index or external system, ranks candidate material, and selects what should enter the model request.

The assembled context may contain:

  • the user’s question;
  • system and developer instructions;
  • relevant conversation state;
  • retrieved source passages and metadata;
  • directions about citation or uncertainty.

The model then generates from its existing parameters and that supplied context. A later layer may check whether claims are supported and whether citations point to the evidence they are supposed to support.

A policy newer than the model

Suppose a model’s documented training cutoff predates a 2026 expense-policy update. A user asks:

Can I expense a monthly coworking pass?

An ungrounded answer would have to rely on parameter knowledge and whatever the question itself implies. A RAG system can create a different path:

2026 policy source
-> retrievable chunks and metadata
-> query for coworking-pass rules
-> selected current-policy passage
-> passage in the model context
-> answer conditioned on that passage

The model can reason over information that did not appear in its training data. Its underlying knowledge cutoff has not moved. Remove the retrieved passage from a later call and the same evidence is no longer available unless another mechanism supplies it again.

That is the distinction between model knowledge and runtime information. RAG extends the runtime information path.

Every stage is a control point

The phrase “give the AI access to our documents” hides the engineering decisions that determine whether useful evidence reaches the model.

Stage Owner Question to inspect
Source coverage Application and source systems Is the current authoritative material present and permitted?
Parsing and chunking Ingestion pipeline Did the retrievable unit preserve the rule and its exceptions?
Query formation Application or retriever Does the search represent the actual information need?
Candidate retrieval Search or retrieval system Did required evidence rank within the candidate set?
Selection and reranking Application policy Which candidates were kept, filtered, or discarded?
Context injection Application or API Did the chosen source text actually enter this call?
Generation Model Did the response use and interpret the evidence correctly?
Grounding and citation check Application, evaluator, or human Do the claims follow from the cited material?

The model owns only one row in that table. It does not ingest the corpus, maintain the index, enforce document permissions, decide which chunks are current, or guarantee that the application supplied the right context.

Retrieval means locating external information for a request. Persistent application memory means retaining information so that it can remain available across interactions. A memory feature may use retrieval to select a past preference, but a search over product documentation is retrieval without being personal conversation memory.

The relationship is:

external information exists
-> retrieval selects candidates
-> application chooses what to supply
-> selected material enters current context
-> model can use it for this call

Calling every external lookup “memory” erases where the information came from and why it was selected. Keep storage policy and retrieval behavior separately inspectable.

RAG helps with freshness and provenance

External sources can be updated without retraining the base model. They can also carry titles, dates, URLs, record IDs, and other provenance that model parameters do not expose precisely.

Those properties make retrieval useful for changing internal policies, product manuals, support records, current code, and other information that needs an inspectable source. A hybrid answer can use model parameters for general explanation while relying on retrieved material for current facts.

Provenance is possible, not automatic. A source label can be wrong. A retrieved passage can be stale. A generated citation can point to a chunk that mentions the topic without supporting the claim. The application must preserve and check the source relationship.

RAG has several independent quality surfaces

A RAG answer can fail even when each component looks plausible in isolation.

  • Source quality: the corpus can contain incorrect, conflicting, outdated, incomplete, or untrusted material.
  • Retrieval quality: search can miss the required passage or rank a merely related passage higher.
  • Selection quality: filtering or reranking can discard the best candidate.
  • Context quality: the application can inject too little, too much, or contradictory evidence.
  • Generation quality: the model can ignore, misread, combine, or overstate what it received.
  • Verification quality: checks can overlook unsupported claims or broken citations.

Retrieved text should also be treated as external input, not automatically trusted instructions. Security controls for hostile content are a wider production topic, but the ownership boundary starts here: the application decides how retrieved material is delimited, filtered, and permitted to influence the run.

Evaluation must follow the stages

A polished answer in one demonstration cannot tell you whether the required evidence was consistently retrieved or whether the model happened to answer from parameter knowledge.

Evaluate at more than one layer:

  1. Test whether known required passages appear in the candidate set.
  2. Inspect ranking and filtering across varied, realistic queries.
  3. Record which source text actually entered the model context.
  4. Check whether answer claims are supported by that text.
  5. Check whether citations identify the supporting source rather than a nearby one.

Benchmarks such as BEIR show why heterogeneous retrieval evaluation matters: performance in one narrow setting can hide weaknesses elsewhere. End-to-end answer checks remain necessary because good candidate retrieval does not force good evidence use.

“Once the documents are in a vector database, RAG makes the model know them and eliminates hallucination.”

The index remains external. For each request, retrieval and application policy select material to place in current context. The model can still receive missing or wrong evidence, misinterpret correct evidence, or generate a claim that the evidence does not support.

The durable mental model

RAG is not one database choice or framework. It is an evidence-delivery architecture:

maintained sources
-> searchable representations
-> candidate retrieval
-> selection and context injection
-> generation
-> support checks

Its value is that model parameters no longer have to be the only information path. Its cost is a new set of application-controlled decisions that must be observed and tested.

The next diagnostic question is therefore not simply “did retrieval find something relevant?” It is “did retrieval find the evidence required for this particular answer?”

References

  1. Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksNeurIPS, 2020
  2. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval ModelsNeurIPS, 2021
  3. Ragas: Automated Evaluation of Retrieval Augmented GenerationShahul Es et al.
  4. Attributed Question Answering: Evaluation and Modeling for Attributed Large Language ModelsBernd Bohnet et al.