Search

How Embeddings Turn Words Into Numbers

The short answer

Quick answer: An embedding is a list of numbers, a vector, that represents a piece of content such as a word, a sentence, an image or a product. Embeddings are produced by a trained model so that things with similar meaning get similar vectors. That turns a fuzzy question, "how alike are these two things?", into simple arithmetic: measure the distance or angle between two vectors. Embeddings are how neural networks take in text, and they power semantic search, recommendations, clustering and retrieval for AI assistants.

Why turn words into numbers

Neural networks calculate with numbers. The naive way to represent words is to give each one an ID, or a one-hot vector: a list as long as the vocabulary, all zeros except a single 1.

That has two problems:

  • It is huge and sparse. A 50,000-word vocabulary needs 50,000 numbers per word.
  • It carries no meaning. "Cat" and "kitten" are exactly as different as "cat" and "carburettor". Every pair of words is equally far apart.

An embedding replaces this with a dense vector of a few hundred to a few thousand numbers, in which position reflects meaning.

The idea: meaning as location

Picture a map where every word is a point. Related words are placed near each other: "dog" near "puppy", "Paris" near "London", "run" near "sprint". An embedding space is that map, but with hundreds of dimensions instead of two.

The individual numbers are not labelled or human-readable. No single dimension means "animal" or "royalty". Meaning lives in the overall pattern and in the relationships between vectors.

Those relationships can be strikingly regular. The best-known example:

vector("king") - vector("man") + vector("woman") ≈ vector("queen")

The direction from "man" to "woman" is roughly the same as from "king" to "queen". Similar directions appear for countries and capitals, or verb tenses. Google's Machine Learning Crash Course on embeddings has a good visual introduction.

How embeddings are learned

Nobody assigns the numbers. They come out of training, guided by a principle from linguistics: you shall know a word by the company it keeps. Words that appear in similar contexts tend to have similar meanings.

Word2vec

The 2013 paper Efficient Estimation of Word Representations in Vector Space made embeddings practical. A small network is trained on a simple task: given a word, predict the words around it (or the reverse).

  1. Start every word with a random vector.
  2. Slide through billions of words of text.
  3. For each word, adjust its vector so it better predicts its neighbours.

"Coffee" and "tea" both appear near "drink", "cup" and "morning", so training pushes their vectors together. The prediction task is discarded afterwards; the vectors are the product. It is ordinary neural network training with an unusual goal.

Contextual embeddings

Word2vec gives each word one fixed vector, which fails for words with several meanings: a river bank and a savings bank get the same one.

Transformers solved this. They start from a fixed embedding for each token, then pass it through attention layers that mix in the surrounding context. The vector that comes out for "bank" differs depending on the sentence.

Sentence and document embeddings

For search and retrieval you want one vector for a whole passage. Embedding models are trained for exactly this, usually with contrastive learning: pairs of texts that should match (a question and its answer, two paraphrases) are pulled together in the space, and unrelated pairs are pushed apart.

Beyond text

The same approach works for images, audio, code, products and users. Models such as CLIP are trained on images paired with captions so that a photo and its description land near each other in a shared space. That is what makes it possible to search photos with a text query, and it underlies AI image generators.

Measuring similarity

Once content is a vector, comparing two items is arithmetic.

MeasureWhat it comparesNotes
Cosine similarityThe angle between two vectorsIgnores length; the most common choice for text. 1 means same direction, 0 unrelated
Dot productAngle and length togetherEquivalent to cosine when vectors are normalised to length 1
Euclidean distanceStraight-line distanceSmaller means more similar
import numpy as np

def cosine_similarity(a, b):
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

Use the measure the embedding model was trained with; its documentation will say which.

What embeddings are used for

  • Semantic search. Embed the query and the documents, and return the nearest documents. A search for "how do I reset my password" finds a page titled "Account recovery steps" with no words in common.
  • Retrieval for AI assistants. Find the passages relevant to a question and give them to a language model. See how RAG works.
  • Recommendations. Represent users and items as vectors and suggest items near a user's tastes. See how recommendation systems work.
  • Clustering and topic discovery. Group similar support tickets or reviews automatically.
  • Classification. Train a small, simple model on top of embeddings.
  • Duplicate and anomaly detection. Near-identical vectors suggest duplicates; isolated ones suggest outliers.
  • Inside language models. The first step of every LLM is to look up an embedding for each token. See how tokenization works.

Searching millions of vectors quickly needs specialised indexes; see what vector databases are.

Practical points

  • Dimensions. More dimensions can capture more nuance but cost more storage and computation. Common sizes range from a few hundred to a few thousand.
  • One model, one space. Vectors from different embedding models are not comparable. If you change model, you must re-embed everything.
  • Chunking. A single vector for a 50-page document blurs its content. Split long texts into passages and embed each.
  • Input limits. Embedding models accept only a certain number of tokens; longer text is truncated.
  • Domain fit. A general-purpose model may do poorly on specialised vocabulary such as legal or medical text. Test on your own data.

Limitations

  • Similar is not the same as correct. "The drug is safe" and "the drug is not safe" are close in the space: same topic, opposite meaning.
  • Bias. Embeddings learn the associations in their training data, including stereotypes.
  • Exact matches. Embeddings are weak at precise identifiers such as part numbers, error codes and names. Combining them with keyword search works better than either alone.
  • Opaque. You cannot read a vector and see why two items are close.

Frequently asked questions

What is an embedding in simple terms?

A list of numbers representing a piece of content, arranged so that similar content has similar numbers.

What is cosine similarity?

A measure of how closely two vectors point in the same direction. It is the standard way of comparing text embeddings.

What is the difference between a token and an embedding?

A token is a piece of text with an ID number. An embedding is the vector of numbers that represents that token, or a longer text, to the model.

How many dimensions does an embedding have?

It depends on the model. Typical values range from a few hundred to a few thousand.

Conclusion

Embeddings translate meaning into geometry. Once words, sentences and images are points in a space where distance reflects similarity, a wide range of problems, from search to recommendation to feeding context to a language model, become questions of finding nearby points.

Related articles

Sources and further reading

Usama Muneer

Usama Muneer

Coder, Blogger, Tech Speaker & Web Technologies Enthusiast. Passionate about working on open-source Programming languages & Tools while utilizing my Product Development skills.

Your experience on this site will be improved by allowing cookies Cookie Policy