CLIOPATRA attacks privacy-preserving LLM insight platforms with malicious chats to leak target medical histories with nearly 100% precision in 65% of cases, showing layered heuristic protections are insufficient.
Regularized LNS turns local search heuristics into MCMC samplers with Fenchel-Young losses, enabling exact block Gibbs sampling and end-to-end learning without global solvers.
LLM-based tree search discovers predictive zebrafish neural models that outperform forecasting baselines, though structural priors are needed to prevent shortcut exploitation and ensure mechanistic recovery.
TokenSwap benchmarks and reduces MLLMs' modality gap by interleaving visual tokens with text, finding reasoning models have smaller gaps and training with TokenSwap mitigates it.
KITScenes Multimodal provides European autonomous driving data with high-fidelity synchronized sensors, complete topologically connected 3D HD maps, and four embodied AI benchmarks.
PaGeR adapts perspective 3D foundation models to panoramas to predict depth, normals, and sky masks in one pass, achieving state-of-the-art 360-degree geometry estimation.
Autonomous diffusion models implicitly optimize a marginal energy landscape via Riemannian gradient flow, where a learned conformal metric neutralizes geometric singularities near data; velocity parameterization ensures stability, but noise prediction fails via a Jensen gap.
AI GameStore proposes evaluating general intelligence via scalable synthesis of human games, finding frontier vision-language models score under 10% of human averages on most generated games.
Croissant Baker generates validated Croissant metadata locally from dataset directories via modular handlers, achieving 97, 100% agreement with ground truth across 140+ datasets including MIMIC-IV.
Project VAANI releases a multimodal dataset of 31,255 speech hours and 289K images spanning 105 Indic languages across 165 Indian districts to support inclusive speech technology.
Echo-GRPO rewrites off-policy reasoning traces into a student VideoLLM's idiolect to avoid gradient clipping, improving reasoning distillation across backbones and benchmarks.
PhysVista benchmarks physical intelligence in vision-language models via a perception-reasoning-assessment loop, exposing major gaps in physical reasoning and plausibility assessment.
BLUE aligns interpretable LLM-generated user profiles with embedding-based recommendation objectives via reinforcement learning, outperforming baselines in sequential recommendation and cross-domain transfer.
Temporal Backtracking Search improves video reasoning by searching over the temporal axis and restarting from verified prefixes rather than resampling from scratch, achieving 22.7% versus 0.7% best-of-N out-of-distribution.
MolmoMotion predicts goal-conditioned 3D point trajectories from visual history and language, outperforming baselines on PointMotionBench and improving robot manipulation and video synthesis.
GenRec separates reconstruction and generation via observation masks to preserve fidelity in visible regions while synthesizing plausible unobserved content.
DualKV eliminates shared-prompt replication in RL training via FlashAttention kernels that process shared and per-sequence KV regions separately, achieving up to 3.82x policy-update speedup.
Self-supervised training on 160,000 in-the-wild videos via a shared coarse mesh yields emergent canonical object frames without pose labels, matching supervised category-level pose estimation accuracy.
A unified framework casts knapsack and top-k operators as dynamic programs with smoothed recursions for differentiable relaxations, parallel algorithms, and theoretical regularization guarantees.
Persona Generators use evolutionary code optimization to expand brief context descriptions into diverse synthetic populations maximizing opinion and preference coverage. Evolved generators substantially outperform baselines across six diversity metrics by spanning rare trait combinations.
Controllable user simulation is formalized as causal inference, proving supervised fine-tuning injects look-ahead bias causing geometric variance explosion and controllability collapse, with proposed mitigations restoring consistency and robust generalization.
The Smart Buildings Control Suite is an open-source HVAC benchmark using multi-year data from 11 buildings and scalable simulators to test control policies across diverse climates and structures.
OpenMHC releases the largest open wearable health dataset with open-source foundation models and a unified benchmark across prediction, imputation, and forecasting tasks.
Differentially private online clustering transforms streams into private semi-coresets via a generic reduction, matching or improving approximation, space, and runtime while inheriting consistency from underlying non-private algorithms.
SPOT-Bench introduces multi-turn proactive queries and Timeliness-F1 to evaluate real-time streaming video perception; AsynKV improves streaming behavior by scaling compute during dead-time to match offline detection.
GlucoFM decomposes CGM data into dual slow and short-term streams for pretraining, improving linear-probe phenotype classification and postprandial response prediction over prior models.
3DSPA is a reference-free 3D semantic point autoencoder that evaluates video realism by integrating point trajectories, depth, and semantic features to detect physical violations and motion artifacts with high human alignment.
HelpBench evaluates LLM privacy, safety, and security advice via 450 questions, finding high average quality but one in ten responses drop below 65% accuracy with harmful errors.
FacePhys is a memory-efficient rPPG algorithm using temporal-spatial state space duality that cuts error by 49% with a 3.6 MB footprint and 9.46 ms latency.
Proposed E-P-R framework diagnoses AI agents consuming conflicting memory via entry-propagation-recovery, finding a compliance trap where early adoption collapses success.
LCDD constructs sparse, causally necessary subnetworks for SFT behaviors, and SFT-Eraser reverses them via activation-matched soft prompts without weight changes.
SkillOS uses RL to train a skill curator that updates an external SkillRepo from experience, improving self-evolving agents across reasoning and multi-turn tasks.
A blackboard multi-agent framework lets autonomous agents volunteer for data-discovery tasks, boosting end-to-end success by 13%-57% over rigid master-slave baselines.
UniVL embeds visual and textual instructions into spatial masks for contextual image generation, cutting FID to 11 and inference costs by 52% without a text encoder.
ACORN strategically selects ML predictions for expert review to optimize occupancy model inference, recovering ecological conclusions near fully human-labeled levels with far fewer reviews.
Orthrus unifies autoregressive and diffusion views in transformers to enable lossless parallel token generation with up to 7.8x speedup and O(1) memory overhead.
ROTATE reformulates ad hoc teamwork as open-ended adversarial training between an agent and teammate generator to probe collaboration deficiencies, substantially improving generalization to unseen partners.
Synthetic noise in pretraining data causes LLM loss divergence with probability scaling by noise type, amount, and model size, exhibiting activation patterns distinct from high-learning-rate failures.
Conditioning diffusion models on multimodal large language models with VAE identity conditioning and dual-layer aggregation improves subject-driven generation by balancing semantics with identity preservation.
DiagnosticIQ benchmarks LLM recommendation of industrial maintenance actions from symbolic rules across 6,690 questions, finding frontier models match human experts but break under structural perturbation due to calibration failures rather than capability gaps.
Uncertainty-guided tree search decouples exploration from policy optimization to bypass RL during exploration, then distills discovered trajectories into deployable policies achieving state-of-the-art sparse-reward results.
PACEvolve++ adapts evolutionary search policies at test time via advisor-model reinforcement learning, using phase-adaptive optimization to outperform frontier-model baselines across engineering and protein tasks.