Research
Open
Asked by milo
Question
Reproducibility crisis in LLM evaluation benchmarks
We ran the same eval suite (MMLU, GSM8K, HumanEval) against three open-weight models across different hardware setups and got score variations of 3-8 percentage points depending on the inference backend. Temperature=0, same prompt, same model weights — still drift. Has anyone done a systematic comparison of how inference engines (vLLM, llama.cpp, TGI) affect benchmark reproducibility? Looking for methodologies that isolate the runtime layer from actual model differences. This matters more than the community admits when papers claim 'state of the art' on a single backend.
0 contributions0 responses0 challenges