LangCode 2027 / Shared Tasks / Task 1

shared_task[1]Multilingual Code Generation

One problem, 100+ natural languages, one correct program. Can your model write the same code no matter which language the problem is posed in?

100+ natural languages Python execution-based Pass@1 macro-avg across tiers
CodaBench · coming soon How it is scored Dates
sum_even.pylive
def sum_even(nums: list[int]) -> int:    """        >>> sum_even([1, 2, 3, 4])    6    """    return sum(n for n in nums if n % 2 == 0)
$ run hidden_tests English ✓ pass

Illustrative problem, not drawn from the test set. Click a language to pin it.

Updates
NAACL 2027LangCode is co-located with NAACL 2027 · June 1–5, 2027 · San Francisco, California, USA Oct 26, 2026First call for papers; shared task training data released Feb 5, 2027Paper submission deadline Mar 12, 2027Commitment deadline for ARR-reviewed papers Mar 26, 2027Notification of acceptance Apr 16, 2027Camera-ready due NewFormal paper submission guidelines are up NewPages for Shared Task 1 and Shared Task 2 are up

##The Task


Code models are prompted mostly in English and benchmarked mostly on English. Task 1 asks how well they hold up when the same programming problem arrives in a different natural language.

01

Read

A function signature and a problem description written in one of more than 100 natural languages, spanning high- to low-resource tiers.

02

Generate

Your system writes a Python solution. One attempt per problem.

03

Execute

The solution runs against hidden unit tests. It passes only if every test passes.

##Evaluation


Systems are ranked by execution-based Pass@1 on hidden tests, macro-averaged across language resource tiers.

Pass@1
The share of problems whose single generated solution passes all hidden tests.
Macro-average
Every resource tier counts equally, so high-resource performance alone cannot win.
Hidden tests
Test cases are never released, which keeps the leaderboard honest.

##Tracks


Participation is free. A constrained track keeps low-compute teams competitive.

Constrained

Small models, big ideas

Limits on model size and compute, so the winning idea does not have to be the biggest GPU budget. Exact limits are announced with the training data.

Open

Anything goes

No limits on model size or compute. Show how far today’s strongest systems reach across languages.

##Data


The task builds on mHumanEval (NAACL 2025), a multilingual extension of HumanEval. The test set uses newly authored, held-out problems to limit contamination from pretraining data, so systems are scored on problems they cannot have memorized.

Released

Training and development data

Problems with prompts across resource tiers, plus public baselines and a starter kit on CodaBench.

Held out

Test problems

New problems with hidden unit tests, released only during the evaluation window.

##Dates


All dates are tentative. Deadlines are 11:59 PM UTC-12 (Anywhere on Earth).

  1. October 26, 2026 Training data releasednext
  2. TBA Evaluation window; dates announced with the training data
  3. February 5, 2027 Paper submission deadline
  4. March 12, 2027 Commitment deadline for ARR-reviewed papers
  5. March 26, 2027 Notification of acceptance
  6. April 16, 2027 Camera-ready due
  7. June 2027 LangCode at NAACL 2027 (June 1–5), San Francisco, California, USA

##Participate


  1. Join the CodaBench competition. The link goes live here with the training data.
  2. Download the training data and starter kit, and build your system.
  3. Submit predictions during the evaluation window and watch the leaderboard.

Questions? Email the organizers at mraihan [at] nd [dot] edu.

Cite the source benchmark

@inproceedings{raihan2025mhumaneval,
  title     = {m{H}uman{E}val - A Multilingual Benchmark to Evaluate Large Language Models for Code Generation},
  author    = {Raihan, Nishat and Anastasopoulos, Antonios and Zampieri, Marcos},
  booktitle = {Proceedings of NAACL 2025},
  year      = {2025},
  url       = {https://aclanthology.org/2025.naacl-long.570/}
}
shared_task[2] → Adversarial Prompt Detection Can a guardrail spot a malicious coding prompt when it is code-mixed or transliterated?