PGLD optimizes synthesis-aware stochastic DNA libraries via policy gradients to bypass synthesis cost limits, enabling million-sequence libraries for antibody exploration at low cost.
SGRPO directly rewards set-level diversity via leave-one-out contributions in a flexible GRPO framework, expanding the utility-diversity Pareto frontier across biomolecular design tasks.
GLACIER treats tandem mass spectrum prediction as graph object detection, outperforming prior state-of-the-art by up to 19.3% on retrieval accuracy with nearly 8-fold faster inference.
EBMol learns atom-additive scalar potentials via flow-inspired restoring field matching to generate physically consistent 3D molecules with state-of-the-art results.
STRATA predicts lipid nanoparticle transfection by aligning molecular structure and composition ratio representations to model component interactions. It improves accuracy and generalizes to unseen molecules and ratios.
PROBE uses edit-response probing to build pocket-specific site maps and EditManuals that guide multi-agent optimization of both affinity and druggability in structure-based drug design, achieving state-of-the-art on CrossDocked2020.
MorphoHELM benchmarks microscopy representation methods across batch effects, finding classic computer vision strategies outperform deep learning across settings and revealing trade-offs between models.
Morph is a flexible-size generative model for 3D molecular design that uses unbalanced optimal transport to dynamically adapt molecular size, improving property steering and enabling out-of-distribution generation.
SKMD introduces symmetry-aware interacting-particle dynamics for active MLIP learning that preserves Boltzmann sampling, yielding faster convergence with fewer training iterations.
CoMole unifies molecular graph generation via motif-aware diffusion and reinforcement learning, achieving top controllability across nine targets with up to 48.2% lower MAE and over 0.94 validity.
Neural compressed sensing extends to function space to co-design wet-lab experiments with learning algorithms, achieving orders-of-magnitude higher information density by measuring multiple molecules simultaneously and deconvolving activity during training for antibodies and cell therapies.
CompleteRXN introduces a benchmark for completing incomplete chemical reaction databases, showing models reach high accuracy on benchmark splits but degrade substantially on uncurated real-world data.
A benchmark of four compositional generalisation tasks reveals state-of-the-art machine learning interatomic potentials fail to generalise to unseen molecules, with out-of-distribution errors often ten times higher than in-distribution errors.
AB-SID-iVAR actively learns Gaussian process targets under unknown self-induced Boltzmann weights, achieving vanishing terminal prediction error without partition function estimation.
Joint Self-Improvement uses a joint generative-predictive model and self-improving sampling to reduce distribution shift and efficiently generate optimized molecules under limited evaluation budgets.
Monroe is a molecular foundation model pre-trained on 81 million molecules that uses in-context TabPFN prediction to achieve state-of-the-art bioassay activity prediction, especially on activity cliffs.
PhenoAIR reformulates Cell Painting mechanism prediction as calibrated evidence reasoning via multi-agent evaluation of noisy retrieved neighbors, outperforming matching and LLM baselines across open-world settings.
ScreenShot is a hierarchical transformer pretrained on drug screening datasets that predicts combination therapy responses via in-context learning from limited observations without molecular profiling or fine-tuning, outperforming baselines and enabling efficient active screening.
Online discrete diffusion adaptation for molecular optimization finds acquisition, reward shaping, and debiasing complementarily boost reward, with replay and validity control stabilizing exploration to outperform offline and search baselines.
Boltz2's protein-ligand co-folding representations match or outperform standalone models on ADMET, generative modeling, and ligand optimization tasks while complementing conventional molecular supervision.