LangCode 2027 / Shared Tasks / Task 2

shared_task[2]Adversarial Prompt Detection

Guardrails trained on English break on other languages. Build one that catches a malicious coding prompt even when it is code-mixed or transliterated.

18 languages ~250K prompts macro-F1 surprise languages
CodaBench · coming soon See the subtasks Dates
guardrail.loglive

Illustrative. Prompt text is redacted; only the programming keywords, which MALICE keeps verbatim, are shown.

Updates
NAACL 2027LangCode is co-located with NAACL 2027 · June 1–5, 2027 · San Francisco, California, USA Oct 26, 2026First call for papers; shared task training data released Feb 5, 2027Paper submission deadline Mar 12, 2027Commitment deadline for ARR-reviewed papers Mar 26, 2027Notification of acceptance Apr 16, 2027Camera-ready due NewFormal paper submission guidelines are up NewPages for Shared Task 1 and Shared Task 2 are up

##The Task


Code-model safety filters are calibrated almost entirely on English. Rewrite a malicious request in a low-resource language, or spell it out in another script, keep the code keywords intact, and many of them stop working.

0.93→0.21 best classifier’s accuracy, English vs. code-mixed malicious prompts
~3× more malicious code generated by even the most safety-aligned model tested
2 styles transliteration into native scripts is nearly as damaging as code-mixing

Findings from MALICE (Raihan, Meher, Dhingra, Zampieri; Findings of EMNLP 2026).

01

Seed

A malicious English coding request from an established security benchmark.

02

Disguise

Translated sentence by sentence into low-resource languages, or transliterated into a non-Latin script, with programming keywords kept verbatim.

03

Detect

Your system decides: benign or adversarial? And if adversarial, what kind of attack?

##Subtasks


Subtask A

Benign or adversarial?

Binary classification of each prompt sent to a code model. This is the main leaderboard.

Subtask B

Which attack?

For adversarial prompts, identify the attack type. Finer-grained, and closer to what a real guardrail needs to log.

Participation is free, and a constrained track keeps low-compute teams competitive.

##Data


The task builds on MALICE, a ~250K-prompt benchmark of code-mixed and transliterated adversarial prompts, split evenly between the two attack styles.

~250K adversarial prompts
18 languages
9 source security benchmarks

Code-mixed 10 languages

Low-resource, Latin script. Sentences are mixed across languages within one prompt.

SwahiliYorubaTagalogSomali HausaIgboCebuanoMalagasy SundaneseJavanese

Transliterated 8 languages

Written in non-Latin scripts (Chinese in Pinyin).

ChinesePinyin Hindiहिन्दी RussianРусский Arabicالعربية Japanese日本語 Korean한국어 GreekΕλληνικά Bengaliবাংলা

Surprise ? languages

Held back until the evaluation window, for a cross-lingual leaderboard.

?????????
Content warning and responsible use. The data contains prompts written to elicit malicious code. Like MALICE itself, it is released for defensive safety research only, under restricted-use terms that participants accept on registration.

##Evaluation


Systems are ranked by macro-F1, so a classifier cannot win by favoring the majority class.

Main leaderboard
Macro-F1 on held-out prompts in the 18 task languages.
Cross-lingual leaderboard
Macro-F1 on surprise languages never seen in training. Does your guardrail generalize?

##Dates


All dates are tentative. Deadlines are 11:59 PM UTC-12 (Anywhere on Earth).

  1. October 26, 2026 Training data releasednext
  2. TBA Evaluation window; surprise languages revealed; dates announced with the training data
  3. February 5, 2027 Paper submission deadline
  4. March 12, 2027 Commitment deadline for ARR-reviewed papers
  5. March 26, 2027 Notification of acceptance
  6. April 16, 2027 Camera-ready due
  7. June 2027 LangCode at NAACL 2027 (June 1–5), San Francisco, California, USA

##Participate


  1. Join the CodaBench competition and accept the research-use terms. The link goes live here with the training data.
  2. Download the training data and starter kit, and build your detector.
  3. Submit predictions during the evaluation window, including on the surprise languages.

Questions? Email the organizers at mraihan [at] nd [dot] edu.

Cite the source benchmark

@inproceedings{raihan2026malice,
  title     = {On the Robustness of Code {LLM} Guardrails to Code-mixed and Transliterated Inputs},
  author    = {Raihan, Nishat and Meher, Dipak and Dhingra, Bhuwan and Zampieri, Marcos},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026}
}
← shared_task[1] Multilingual Code Generation One problem, 100+ natural languages, one correct program.