Adding response times to preference data via drift-diffusion modeling restores identifiability of average preferences among anonymous heterogeneous labelers, correcting choice-only estimation bias without tracking users.
Gated DeltaNet-2 decouples erase and write with channel-wise gates to improve linear attention, achieving top performance among recurrent and hybrid models at 1.3B scale with strong long-context retrieval.
Interactive video world models lose visual persistence beyond training horizons because temporal RoPE offsets become out-of-distribution; WorldTrace assigns compressed memory slots virtual in-distribution positions to restore addressability, boosting temporal consistency by 15.5% and episodic recall
RigidFormer is a transformer that learns mesh-free rigid-body dynamics via object-level anchors and differentiable Kabsch projection, outperforming mesh-based baselines with faster inference and scalability to 200+ objects.
Time-shifted anechoic targets, a two-stage distortion-perception framework, and curated data improve universal speech enhancement and achieve state-of-the-art results.
RelationVGGT enables feed-forward 3D spatial relation segmentation across multi-view images without camera poses or category names by combining visual semantics with geometry-aware representations via a relation transformer.
STRABLE introduces 108 real-world string-and-number tables and benchmarks 445 pipelines, finding simple embeddings with advanced learners suffice for categorical tables while LLMs help on free-text tables.
MulTaBench benchmarks 40 multimodal tabular datasets and shows target-aware tuning of text and image embeddings improves predictive performance over frozen embeddings.