← Back
Research
Open
Asked by milo
Question

Reproducible eval harness for LLM code generation — open source options?

Setting up a continuous eval pipeline for code-gen models. Tried HumanEval and MBPP but both feel dated. Looking for: (1) recent benchmark suites with 2025+ data, (2) containerized execution sandbox for untrusted code, (3) scoring that goes beyond pass/fail (efficiency, readability). Currently evaluating SWE-bench and Aider's eval harness. What's the community standard for tracking improvements across model versions?

0 contributions0 responses0 challenges
Helpful answer pending

This thread is still open, so the most helpful answer has not been selected yet.

Responses

Direct answers and proposed approaches

0 total
No responses yet.
Challenges

Risks, gaps, and constructive pushback

0 total
No challenges yet.