##The Task
Code models are prompted mostly in English and benchmarked mostly on English. Task 1 asks how well they hold up when the same programming problem arrives in a different natural language.
Read
A function signature and a problem description written in one of more than 100 natural languages, spanning high- to low-resource tiers.
Generate
Your system writes a Python solution. One attempt per problem.
Execute
The solution runs against hidden unit tests. It passes only if every test passes.
##Evaluation
Systems are ranked by execution-based Pass@1 on hidden tests, macro-averaged across language resource tiers.
- Pass@1
- The share of problems whose single generated solution passes all hidden tests.
- Macro-average
- Every resource tier counts equally, so high-resource performance alone cannot win.
- Hidden tests
- Test cases are never released, which keeps the leaderboard honest.
##Tracks
Participation is free. A constrained track keeps low-compute teams competitive.
Small models, big ideas
Limits on model size and compute, so the winning idea does not have to be the biggest GPU budget. Exact limits are announced with the training data.
Anything goes
No limits on model size or compute. Show how far today’s strongest systems reach across languages.
##Data
The task builds on mHumanEval (NAACL 2025), a multilingual extension of HumanEval. The test set uses newly authored, held-out problems to limit contamination from pretraining data, so systems are scored on problems they cannot have memorized.
Training and development data
Problems with prompts across resource tiers, plus public baselines and a starter kit on CodaBench.
Test problems
New problems with hidden unit tests, released only during the evaluation window.
##Dates
All dates are tentative. Deadlines are 11:59 PM UTC-12 (Anywhere on Earth).
- October 26, 2026 Training data releasednext
- TBA Evaluation window; dates announced with the training data
- February 5, 2027 Paper submission deadline
- March 12, 2027 Commitment deadline for ARR-reviewed papers
- March 26, 2027 Notification of acceptance
- April 16, 2027 Camera-ready due
- June 2027 LangCode at NAACL 2027 (June 1–5), San Francisco, California, USA
##Participate
- Join the CodaBench competition. The link goes live here with the training data.
- Download the training data and starter kit, and build your system.
- Submit predictions during the evaluation window and watch the leaderboard.
Questions? Email the organizers at mraihan [at] nd [dot] edu.
Cite the source benchmark
@inproceedings{raihan2025mhumaneval,
title = {m{H}uman{E}val - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation},
author = {Raihan, Nishat and Anastasopoulos, Antonios and Zampieri, Marcos},
booktitle = {Proceedings of NAACL 2025},
year = {2025},
url = {https://aclanthology.org/2025.naacl-long.570/}
}