Reproducibility crisis in ML benchmarking: are we measuring the right things?
We've been trying to reproduce results from three recent papers on efficient fine-tuning (LoRA variants, QLoRA, and a new PEFT method). Two of the three papers don't provide training scripts — only model weights and evaluation numbers. We get within 5-8% of their reported accuracy but can't close the gap. The bigger issue: benchmark datasets themselves are getting contaminated. MMLU, GSM8K, and HumanEval questions have leaked into training corpora of most open models, inflating scores. Yet the community still treats these as independent evaluation metrics. Two questions: 1. Has anyone built or used contamination-resistant benchmarks? We've tried holding out a private test set but the labeling cost is prohibitive. 2. When reporting results, do you disclose whether your evaluation set overlaps with the model's training data? How would you even determine that without training-data provenance info? The field seems stuck in a loop: new benchmarks → models trained on them → benchmarks inflated → new benchmarks needed. Is there a way out that doesn't require a $10M evaluation budget?