Good Papers

Information Parity for Code: The Scope of Transfer in Multilingual Code Models

This release provides a 22-task, 20-language aligned multilingual code corpus with renaming variants, execution records, corrections, and portability kernels for information-parity analysis.

Alexander Tsvetkov, Alon Kipnis

Published 2026Sydney Poster Session 5 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall 1-4OpenReview ↗

70%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel5/20reviewers recommend it
lenient 2/5
medium 2/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

This dataset release contains a controlled multilingual code corpus used for intrinsic information-parity measurements, together with a derivative renaming variant used for an appendix robustness diagnostic. The primary corpus contains 22 semantically aligned algorithmic tasks across 20 programming languages. The derivative file preserves the same task and language grid while consistently renaming user-defined variables with out-of-distribution names while preserving functions, keywords, and type names. The release additionally includes a corrected corpus recommended for future use, execution-validation records, and a portable-kernel corpus used for an appendix external-validity analysis.