Fundamentals of Generative AI
Embeddings and Vector Search
Embeddings represent content as numeric vectors that capture meaning. Vector search compares them to find related content even when exact words differ, which is central to semantic search and RAG.
Concepts
- Embeddings are numeric vectors that represent semantic meaning for text, images, or other data.
- Vector databases and vector indexes support similarity search for embeddings.
- Semantic search finds meaningfully related content even when exact keywords differ.
- Similar concepts have nearby embeddings even when their wording differs.
- Similarity search ranks embeddings by semantic closeness rather than exact keyword matches.
- Amazon Bedrock Knowledge Bases supports vector stores including Amazon OpenSearch Serverless, Amazon Aurora PostgreSQL with pgvector, Pinecone, and Redis Enterprise Cloud.
Exam tips
- Embeddings are central to semantic search and RAG retrieval.
- Vector indexes or vector databases support similarity search over embeddings.
- Choose semantic search when meaning matters more than exact keyword matching.
- Embeddings enable meaning-based retrieval when queries and sources use different words.
- Amazon OpenSearch Serverless is the default and most common vector store for Amazon Bedrock Knowledge Bases.
- Bedrock Knowledge Bases can also use Amazon Aurora PostgreSQL with pgvector, Pinecone, and Redis Enterprise Cloud as supported vector stores.