How Meaning Becomes Math: Embeddings

Embeddings turn text into numeric vectors whose relative positions can rank related candidates; that ranking enables semantic search without proving that a result is true or sufficient.

  • Explainer
  • 6 min read
Illustration showing text mapped into vector space coordinates for semantic search.

A search for end membership should probably find a document titled Cancel a subscription, even though the words do not match. Keyword search can be extended with synonyms and other techniques, but embedding search offers a general mechanism for comparing learned relatedness.

An embedding is a numeric vector representation of an item such as a passage, image, or query. This lesson focuses on text. A model maps text to a list of numbers, and a retrieval system compares those lists so that more related representations can rank nearer one another.

The useful mental model is semantic coordinates. It is an analogy, not a literal two-dimensional map and not a claim that every coordinate has a readable meaning.

One illustrative similarity search

Values and ranking are explanatory, not measured
  1. Text / 01"Can I end my plan?"

    A query begins as text.

  2. Embedding model / 02[0.14, -0.62, ..., 0.09]

    The model maps it to a numeric vector.

  3. Comparison / 03Rank nearby vectors

    A similarity function compares the query with indexed candidates.

Candidate order

  1. 01 / closerCancel a subscriptionRelated intent
  2. 02 / nearbyPause an accountRelated, not identical
  3. 03 / fartherAdd a tax number to an invoiceDifferent task
What the ranking establishesCandidate relatedness, not truth, completeness, or task sufficiency
Embeddings make relatedness rankable.Both queries and indexed items receive model-specific vector representations. Applications compare those vectors to rank candidates, then retrieve the source material associated with selected vectors. The vector is not a lossless copy of the text.

From text to a ranking signal

Imagine a help center with thousands of passages. A basic embedding retrieval path looks like this:

  1. Split source documents into retrievable passages, usually called chunks.
  2. Run each chunk through an embedding model to produce a vector.
  3. Store each vector with a reference to its source text and useful metadata.
  4. Embed the user’s query with a compatible model.
  5. Compare the query vector with indexed vectors using a similarity or distance function.
  6. Return the highest-ranked candidates for later selection or use.

Cosine similarity is one common comparison. Other metrics and index strategies exist, and their behavior depends on the representation and implementation. A beginner does not need the formula to understand the causal role: the comparison converts relative vector positions into an ordering.

That ordering is the output of similarity search. It is not yet an answer.

What the vector does and does not preserve

Calling an embedding a compressed document creates the wrong expectation. A typical retrieval system does not reconstruct the original passage from its embedding. It uses the vector as a comparison representation, then follows the stored association back to the source chunk.

The vector can preserve relationships that are useful for the model’s training objective, but it is not a lossless container for every fact, phrase, or distinction in the text. Different embedding models can represent the same passage differently. A model can also blur distinctions that matter for a particular task.

So keep three objects separate:

Object Role
Source chunk The actual text that may contain evidence
Embedding vector A numeric representation used for comparison
Similarity score or rank A signal about relative relatedness under this setup

The source chunk may be useful evidence. The vector and score help find it.

Consider these phrases:

cancel a subscription
end my membership
close the paid plan
add a tax number to an invoice

A useful embedding model may place the first three in a related region of its representation space even though they share few exact words. The fourth should usually rank farther away for a cancellation query.

This capability is why embeddings support semantic retrieval: matching can follow learned relationships rather than exact token overlap alone.

There are limits. pause my account may be close to cancellation while expressing a different operation. contractor access may be close to employee access while the category difference decides the policy answer. Similarity can recover a neighborhood and still miss the required address.

Candidate is the precise word

Suppose an internal handbook receives the query:

Can contractors use single sign-on?

The search might rank these chunks:

Candidate Why it looks related What remains unresolved
Employees sign in with SSO Authentication and SSO are close to the query It says nothing about contractors
Contractors require temporary identity approval Contractor identity is close to the query It may contain part of the required rule
SSO outage response It shares the exact term SSO It addresses availability, not eligibility

All three can be legitimate candidates. None becomes correct merely by ranking first. The application still has to decide which material is useful, retrieve the associated source text, place selected evidence into context, and assess whether an answer is supported.

This is why nearest neighbor means a nearby vector under the chosen comparison, not “the passage that must contain the answer.”

Chunking changes the searchable unit

Embedding search usually operates on chunks rather than whole document libraries. Chunk boundaries therefore change what the system can retrieve.

Imagine a policy paragraph:

Refund requests close after 30 days.
When a fraud review remains open, the deadline begins after the review closes.

If preprocessing separates the exception from the rule, a query about the fraud-review deadline may retrieve the generic first sentence without the condition that changes the answer. A very large chunk can create a different problem: the relevant detail may be surrounded by enough unrelated material that its representation becomes less useful for the query.

Overlap can keep nearby conditions together across chunk boundaries, but more overlap is not automatically better. It can duplicate candidates, consume storage and context, and still fail when the underlying structure is poorly represented. Useful chunk size and overlap depend on document structure, query patterns, model behavior, and evaluation results.

Chunking is therefore not clerical preprocessing. It defines the units that retrieval can see.

Similarity is not truth

An embedding system can rank a false statement close to a query because the statement is directly about the topic. It can rank an outdated policy above a current one. It can retrieve a passage that names the right entity but omits a decisive date or exception.

Similarity does not establish:

  • whether the source is authoritative;
  • whether its facts are current;
  • whether it contains every necessary condition;
  • whether it addresses the user’s exact entity or jurisdiction;
  • whether a later model will interpret it correctly.

Those are downstream questions about source quality, filtering, evidence fit, context assembly, and generation.

“An embedding contains the meaning of a document, so the closest vector is the answer.”

An embedding is a model-specific representation used to compare items. Similarity produces a ranked candidate set. The system must still retrieve source material and determine whether it is adequate evidence for the task.

Evaluate the retrieval behavior you need

A successful demonstration proves that one query worked in one corpus. It does not establish that the same setup is robust across topics, writing styles, languages, document structures, or unfamiliar query types.

The BEIR benchmark was created in response to retrieval models being studied in narrow, homogeneous settings. Its broader lesson applies beyond any one benchmark: test retrieval across the variety the real system will face.

Useful evaluation starts with concrete questions:

  • Did the required chunk appear among the returned candidates?
  • How often are the top results irrelevant or incomplete?
  • Do filters preserve the right date, product, tenant, or jurisdiction?
  • Which query types fail after changing the embedding model or chunking policy?

Embeddings make semantic candidates discoverable. They do not decide which candidate becomes context or what the model says next. That larger assembly is the job of a retrieval-augmented generation pipeline.

References

  1. Embeddings guideOpenAI
  2. BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval ModelsNeurIPS, 2021
  3. Chunk large documents for RAG and vector search in Azure AI SearchMicrosoft Learn