- Published:
- Last updated:
Embeddings, Vector Search, and RAG
- Authors
- Name
- Lucian Oprea
- @LucianDSA_
00:30:00
Filter by difficulty
33 of 33 questions shown
RAG and Embedding Foundations
⏷ 1. What is retrieval-augmented generation, and what problem does it solve?
⏷ 2. What is an embedding?
⏷ 3. How do the offline indexing and online query pipelines work?
⏷ 4. How do cosine similarity, dot product, and Euclidean distance differ?
⏷ 5. Why split documents into chunks instead of embedding the whole document?
Free preview complete
You’ve reached the end of the free preview
Get every remaining question and complete answer, plus progress tracking across the full Interview Question Library.
- 28 more questions and complete answers in this topic
- Full access to every interview topic
- Progress tracking and question flags
- New questions and improvements during your subscription
Full access from
$12/month
No long-term commitment. Cancel whenever you want.
⏷ 6. How do you choose an embedding model and vector dimension?
⏷ 7. Why must queries and documents use compatible embeddings?
⏷ 8. Which changes require re-embedding or reindexing?
⏷ 9. How do dense, sparse, and multimodal embeddings differ?
Vector Indexes and Search Infrastructure
⏷ 10. What is the difference between exact kNN and approximate nearest-neighbor search?
⏷ 11. How does an HNSW index work, and what are its tradeoffs?
⏷ 12. How do IVF-style indexes work, and what are their tradeoffs?
⏷ 13. How do you tune recall, latency, memory, and cost for vector search?
⏷ 14. How do metadata filters interact with ANN search?
⏷ 15. How do you choose between a vector database, search engine, and relational vector extension?
Ingestion, Chunking, and Freshness
⏷ 16. How do you design an idempotent document-ingestion pipeline?
⏷ 17. How do you choose a chunking strategy?
⏷ 18. When does chunk overlap help, and when does it hurt?
⏷ 19. How do you preserve tables, code, layout, and provenance during chunking?
⏷ 20. Which metadata should a RAG chunk store?
⏷ 21. How do you handle document updates, freshness, and deletion?
Retrieval, Reranking, and Context Assembly
⏷ 22. When should you use keyword, semantic, or hybrid retrieval?
⏷ 23. How do you combine lexical and vector rankings?
⏷ 24. When do query rewriting, decomposition, and multi-query retrieval help?
⏷ 25. What does a reranker do, and when is it worth the cost?
⏷ 26. How do you assemble a useful context within a token budget?
⏷ 27. How should citations and grounding work in a RAG answer?
Evaluation, Security, and Scale
⏷ 28. How do you evaluate retrieval with Recall@k, Precision@k, MRR, and nDCG?
⏷ 29. Why must you evaluate retrieval and answer generation separately?
⏷ 30. How do you secure a multi-tenant RAG system end to end?
Advanced Retrieval Patterns
⏷ 31. How does parent-child retrieval balance precise search with coherent context?
⏷ 32. How do single-vector retrieval, late interaction, and cross-encoder reranking differ?
⏷ 33. When should a RAG system route, change strategy, or skip retrieval?