Optimize generative AI systems and model performance
Retrieval-Augmented Generation Optimization
CoreTune chunking, similarity thresholds, and top-k retrieval, choose and size embedding models, combine keyword and vector search with semantic ranking, and measure retrieval quality before and after changes.
Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.
Why this matters
Most RAG failures are retrieval failures: the right passage was never found or was buried under noise. Tuning chunks, embeddings, and hybrid ranking, then measuring the effect, is the most direct way to improve grounded answers.
Must Know
- Microsoft recommends starting with chunks of about 512 tokens with roughly 25 percent overlap, then tuning. Smaller chunks give precise matches with less context, and larger chunks keep context but dilute relevance.
- The Text Split skill chunks by pages or sentences with a configurable maximum length and overlap. Structure-aware chunking with document layout keeps headings and sections together.
- Hybrid search runs keyword (BM25) and vector queries in parallel and merges results with Reciprocal Rank Fusion. Semantic ranking then reranks the top 50 merged results.
- Raising a similarity or relevance threshold, or lowering top-k, removes weak matches but can drop useful context; lowering it does the opposite.
- text-embedding-3-large produces 3,072 dimensions and text-embedding-3-small 1,536; both accept a dimensions parameter to shorten vectors and reduce storage.
- The query-time vectorizer must use the same embedding model as the indexed content, so changing the embedding model means re-embedding the corpus.
- HNSW performs approximate nearest neighbor search for speed, while exhaustive KNN finds the exact nearest neighbors at higher cost.
- Document retrieval evaluation reports metrics such as NDCG, fidelity, and holes against relevance labels, and comparisons between configurations should use the same queries.
Compare and Distinguish
- Keyword versus vector versus hybrid retrieval: exact terms versus meaning versus both, merged with RRF.
- Hybrid ranking versus semantic ranking: score fusion versus a language-model reranker on the top results.
- Smaller versus larger chunks: precision versus context.
- HNSW versus exhaustive KNN: approximate and fast versus exact and expensive.
Scenario examples
- Scenario: Searches for exact product codes fail with vector-only retrieval. Think: hybrid search so keyword matching catches the codes.
- Scenario: Answers miss details split across chunk boundaries. Think: add chunk overlap or use structure-aware chunking.
- Scenario: A team switches to a new embedding model. Think: re-embed every document and update the query vectorizer together.
Exam traps
- Vectors from different embedding models are not comparable, even at the same dimension count.
- Semantic ranking reorders results but does not retrieve documents that keyword and vector search missed.
- Raising top-k adds context but also adds noise and tokens.
- Text Split overlap must be less than half the maximum page length.
Key takeaways
- Tune chunking and thresholds with measured retrieval metrics.
- Use hybrid search with semantic ranking for mixed keyword and meaning queries.
- Treat embedding model changes as full reindexing events.
How it works
- RRF scores each document by summing 1/(rank + k) across result lists, rewarding documents ranked well in several lists.
- Integrated vectorization chunks and embeds content during indexing and vectorizes queries at search time.
Objects and administrative surfaces
- Azure AI Search indexes, vector profiles, skillsets, and semantic configurations.
- Integrated vectorization with embedding skills and query-time vectorizers.
- Foundry retrieval evaluators for comparing configurations.
When to use it
- Use exhaustive KNN for small indexes or to measure HNSW recall, and HNSW for large production indexes.
- Use shorter embedding dimensions when storage or latency matters more than the last points of accuracy.
Security and governance implications
- Apply document-level security trimming so retrieval returns only content the user may see.
- Track index and embedding configuration changes in source control.
Troubleshooting signals
- Low groundedness with relevant documents in the index often points to chunks that split key facts.
- Many irrelevant passages in context suggest a threshold that is too low or a top-k that is too high.
More detail
- Tune similarity thresholds, chunk sizes, and retrieval strategies.
- Select and adapt embedding models for domain accuracy.
- Implement hybrid search with semantic ranking.
- Evaluate retrieval with relevance metrics and controlled comparisons.
Ready for the quiz?
- How does Reciprocal Rank Fusion combine keyword and vector results?
- What is the recommended starting chunk size and overlap?
- Why must the query vectorizer match the index embedding model?
- What does the dimensions parameter trade off?
Related objectives
- D5.1.S1 — Optimize retrieval performance by tuning similarity thresholds, chunk sizes, and retrieval strategies
- D5.1.S2 — Select and fine-tune embedding models for domain-specific use cases and accuracy improvements
- D5.1.S3 — Implement and optimize hybrid search approaches combining semantic and keyword-based retrieval
- D5.1.S4 — Evaluate and improve RAG system performance by using relevance metrics and A/B testing frameworks