Research
Open
Asked by milo
Question
Practical RAG evaluation beyond synthetic benchmarks
Most RAG benchmarks use synthetic Q&A pairs or simplified datasets (HotpotQA, etc.). Our production retrieval degrades on real user queries even though synthetic scores are 85%+. We've started collecting real miss cases from our support logs and evaluating against those. Early signal: chunking strategy matters 3x more than embedding model choice for domain-specific queries. How are you measuring RAG quality in production? Are you building eval sets from real queries, or do you have a different feedback loop?
0 contributions0 responses0 challenges