Good Papers

mRNABench: A curated benchmark for mature mRNA property and function prediction

mRNABench benchmarks mature mRNA property predictions across 59 tasks and 135K experiments, revealing synergies between self-supervised objectives that yield a compact state-of-the-art Mamba model using 700x fewer parameters.

Ruian (Ian) Shi, Taykhoom Dalal, Philip Fradkin, Divya Koyyalagunta, Simran Chhabria, Andrew Jung, Cyrus L Tam, Defne Ceyhan, Jessica Lin, Kaitlin U Laverty, Ilyes Baali, Bo Wang, Quaid Morris

Published 2026Sydney Poster Session 2 · Tue, Dec 8, 5:00 PM–8:00 PM local time · Hall 1-4OpenReview ↗

91%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel18/20reviewers recommend it
lenient 5/5
medium 9/10
strict 4/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Messenger RNA (mRNA) is central in gene expression, and its half-life, localization, and translation efficiency drive phenotypic diversity in eukaryotic cells. While supervised learning has widely been used to study the mRNA regulatory code, self-supervised foundation models support a wider range of transfer learning tasks. However, the dearth and homogeneity of standardized benchmarks limit efforts to pinpoint the strengths of various models. Here, we present mRNABench, a comprehensive benchmarking suite for mature mRNA biology that evaluates the representational quality of mature mRNA embeddings from self-supervised nucleotide foundation models. We curate ten datasets and 59 prediction tasks that broadly capture salient properties of mature mRNA, and assess the performance of 18 families of nucleotide foundation models for a total of 135K experiments. Using these experiments, we study parameter scaling, compositional generalization from learned biological features, and correlations between sequence compressibility and performance. We identify synergies between two self-supervised learning objectives, and pre-train a new Mamba-based model that achieves state-of-the-art performance using 700x fewer parameters. mRNABench can be found at: https://github.com/morrislab/mRNABench.