Aleph Alpha changes benchmark filters after finding missed code solutions
New analysis explains why copied solutions escaped checks and how exposure changes proxy benchmark scores.

Aleph Alpha detailed gaps in its benchmark contamination checks on October 6, explaining how copied code solutions escaped filters that expected matching questions.
The company says it now indexes HumanEval and MBPP reference solutions separately. For four HumanEval variants with fixed answer strings, it now indexes questions alone. Matches require review because short solutions can be natural implementations.
In a smaller model experiment, restoring documents at expected production exposure raised HumanEval by 19 percentage points and MMLU by seven against a filtered baseline. These are proxy results, not Kolibri performance gains.
Its earlier technical report already warned that Kolibri could repeat HumanEval solutions learned during pretraining. The new analysis says filtering must extend to that earlier data. Paraphrases and translations remain blind spots.
Illustrative archival coding photograph by Goran Ivos under CC0. Resized and converted to WebP. The image does not show the experiments.



