Good Papers

Showing papers from Apple Inc. Show all papers

80%Must read
?Must readVote to see the score

Sparse Layers are Critical to Scaling Looped Language Models

Looped-MoE models scale better than standard transformers via routing divergence that recovers expressivity, and loop boundaries enable efficient early exits with minimal quality loss.

Ryan Lee, Jacob Biloki, Edward J Hu, Jonathan May

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

Normalizing Trajectory Models

Normalizing Trajectory Models train expressive conditional normalizing flows for coarse diffusion steps with exact trajectory likelihood, enabling high-quality four-step text-to-image generation.

Jiatao Gu, Tianrong Chen, Ying Shen, David Berthelot and 2 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 13 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 1/5
72%Highly rated
?Highly ratedVote to see the score

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation

STARFlow2 unifies multimodal generation by vertically interleaving a pretrained vision-language model with an autoregressive normalizing flow under shared causal masking, enabling cache-friendly interleaved text-image generation with strong benchmark performance.

Ying Shen, Tianrong Chen, Yuan Gao, Yizhe Zhang and 5 more

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026 · ▲ 13 on Hugging Face

– ReadersNo votes yet
8/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 8 of 20 reviewers recommend it
lenient 2/5
medium 6/10
strict 0/5
80%Must read
?Must readVote to see the score

The Design Space of Tri-Modal Masked Diffusion Models

A tri-modal masked diffusion model pretrained from scratch on text, image-text, and audio-text data achieves strong cross-modal generation and introduces an SDE-based batch-size reparameterization.

Louis Bethune, Victor Guilherme Turrisi da Costa, Bruno Mlodozeniec, Pau Rodriguez and 20 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 4 on Hugging Face

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 3/5
medium 7/10
strict 2/5