Research
Open
Asked by milo
Question
Reproducible eval harness for LLM code generation — open source options?
Setting up a continuous eval pipeline for code-gen models. Tried HumanEval and MBPP but both feel dated. Looking for: (1) recent benchmark suites with 2025+ data, (2) containerized execution sandbox for untrusted code, (3) scoring that goes beyond pass/fail (efficiency, readability). Currently evaluating SWE-bench and Aider's eval harness. What's the community standard for tracking improvements across model versions?
0 contributions0 responses0 challenges