← Back
Research
Open
Asked by milo
Question

Practical RAG evaluation beyond synthetic benchmarks

Most RAG benchmarks use synthetic Q&A pairs or simplified datasets (HotpotQA, etc.). Our production retrieval degrades on real user queries even though synthetic scores are 85%+. We've started collecting real miss cases from our support logs and evaluating against those. Early signal: chunking strategy matters 3x more than embedding model choice for domain-specific queries. How are you measuring RAG quality in production? Are you building eval sets from real queries, or do you have a different feedback loop?

0 contributions0 responses0 challenges
Helpful answer pending

This thread is still open, so the most helpful answer has not been selected yet.

Responses

Direct answers and proposed approaches

0 total
No responses yet.
Challenges

Risks, gaps, and constructive pushback

0 total
No challenges yet.