Adding response times to preference data via drift-diffusion modeling restores identifiability of average preferences among anonymous heterogeneous labelers, correcting choice-only estimation bias without tracking users.
Gated DeltaNet-2 decouples erase and write with channel-wise gates to improve linear attention, achieving top performance among recurrent and hybrid models at 1.3B scale with strong long-context retrieval.
Interactive video world models lose visual persistence beyond training horizons because temporal RoPE offsets become out-of-distribution; WorldTrace assigns compressed memory slots virtual in-distribution positions to restore addressability, boosting temporal consistency by 15.5% and episodic recall
RigidFormer is a transformer that learns mesh-free rigid-body dynamics via object-level anchors and differentiable Kabsch projection, outperforming mesh-based baselines with faster inference and scalability to 200+ objects.
Time-shifted anechoic targets, a two-stage distortion-perception framework, and curated data improve universal speech enhancement and achieve state-of-the-art results.
RelationVGGT enables feed-forward 3D spatial relation segmentation across multi-view images without camera poses or category names by combining visual semantics with geometry-aware representations via a relation transformer.
STRABLE introduces 108 real-world string-and-number tables and benchmarks 445 pipelines, finding simple embeddings with advanced learners suffice for categorical tables while LLMs help on free-text tables.
MulTaBench benchmarks 40 multimodal tabular datasets and shows target-aware tuning of text and image embeddings improves predictive performance over frozen embeddings.
Sparse Koopman autoencoders use sparse latent supports as label-free regime indicators that identify local dynamical basins and outperform dense autoencoders in multibasin forecasting.
Physical AI Smart Spaces introduces multi-camera 3D perception benchmarks for indoor smart spaces spanning synthetic and real-world data, plus 3D HOTA evaluation.
Grounded-Exo2Ego couples geometric anchoring with semantic grounding and camera relocalization to robustly generate egocentric video from exocentric inputs, outperforming prior methods on EgoExo4D.
RePlaid, a continuous diffusion language model aligned with modern discrete architectures, achieves scaling laws rivaling discrete diffusion and sets a continuous diffusion perplexity record of 22.1 on OpenWebText.
CDM amortizes twisted SMC for discrete diffusion by learning a twist function via contrastive samples, adding under 5% overhead while outperforming baselines on text, DNA, protein, and LLM tasks.
DiLaDiff proposes a latent-augmented masked diffusion language model with consistency distillation that improves quality and accelerates inference by generating continuous latents in negligible time.
Flow map denoisers implicitly define a one-parameter family spanning the distortion-perception tradeoff via lookahead parameter t, matching or exceeding specialized baselines across inverse problems.
GeoTransolver extends Transolver with geometry-aware attention and multi-scale ball queries to improve operator learning accuracy and robustness on irregular engineering domains.
PiD reformulates latent decoding as conditional pixel diffusion to synthesize high-resolution images with low latency and high fidelity. It decodes 512×512 latents to 2048×2048 pixels in under one second on consumer GPUs.
PointZero predicts full 3D point tracks from sparse tracks and RGB-D to learn transferable dynamics without robot actions, outperforming baselines on dynamics and manipulation tasks.
SKMD introduces symmetry-aware interacting-particle dynamics for active MLIP learning that preserves Boltzmann sampling, yielding faster convergence with fewer training iterations.
Flash-KMeans eliminates GPU HBM bottlenecks via fused assignment and inverse mapping updates, delivering up to 17.9x speedups over existing exact k-means implementations.
MARS introduces adaptive rank search balancing multimodal convergence dynamics via dual scaling laws to optimize low-rank fine-tuning of multimodal large language models.
FedRevive revives stale asynchronous federated updates via server-side data-free knowledge distillation, accelerating training by 38.4% and boosting accuracy by 16.5%.
LB-MCTS combines tree-structured search with Bayesian optimization and large language models to solve CASH, outperforming baselines on 104 datasets via adaptive reliability-aware proposal shifts.
DuplexPO decouples conversational dynamics from reasoning via RL, improving turn-taking and backchannels without sacrificing instruction-following ability.
Spatial-IQ hierarchically decomposes spatial reasoning into perceptual and cognitive sub-tasks, showing models use shortcuts and that hierarchical chain-of-thought training improves consistency and accuracy.
SpatialClaw uses a stateful Python kernel with step-wise code execution to enable flexible spatial reasoning, achieving 59.9% average accuracy across 20 benchmarks.
A reasoning-prefix masking framework distills think-answer visual reasoning into compact VLMs by masking salient reasoning cues to force visual anchoring, improving multimodal benchmarks over prior distillation methods.