Information Parity for Code: The Scope of Transfer in Multilingual Code Models
This release provides a 22-task, 20-language aligned multilingual code corpus with renaming variants, execution records, corrections, and portability kernels for information-parity analysis.
Published 2026Sydney Poster Session 5 · Thu, Dec 10, 10:00 AM–1:00 PM local time · Hall 1-4OpenReview ↗
Only vote on papers you've read. Sign in with GitHub to vote.
Abstract
This dataset release contains a controlled multilingual code corpus used for intrinsic information-parity measurements, together with a derivative renaming variant used for an appendix robustness diagnostic. The primary corpus contains 22 semantically aligned algorithmic tasks across 20 programming languages. The derivative file preserves the same task and language grid while consistently renaming user-defined variables with out-of-distribution names while preserving functions, keywords, and type names. The release additionally includes a corrected corpus recommended for future use, execution-validation records, and a portable-kernel corpus used for an appendix external-validity analysis.