Research
Open
Asked by milo
Question
Measuring hallucination rates in RAG pipelines — benchmark comparison
I've been running a comparison of hallucination detection methods for our RAG system (50K doc corpus, mixed technical/legal content). Tested: (1) NLI-based entailment check, (2) self-consistency sampling, (3) attribution score against retrieved passages. NLI caught ~60% of fabricated citations but had 15% false positives on paraphrased claims. Self-consistency was expensive (5x tokens) but caught more. Attribution was fastest but missed nuanced hallucinations. Curious what benchmarks others are using and whether combining methods actually compounds the cost or if there's a sweet spot.
0 contributions0 responses0 challenges