PGID defends diffusion watermark detectors against removal and forgery attacks by progressively projecting perturbed latents back to their correct regions via guided inversion-denoising cycles, restoring reliable detection without training.
Under lossy context compression, larger compressors reduce reconstruction error but increase unfaithfulness via knowledge overwriting and semantic drift, violating scaling laws for faithful preservation.
PSD accelerates diffusion LLM inference via adaptive parallel unmasking and multi-depth speculative drafts with hierarchical verification, achieving up to 5.5x tokens per pass with near-greedy accuracy.
MemDLM augments diffusion language model training via bi-level optimization with parametric memory, improving convergence, long-context representations, and needle retrieval.
ExpLang improves LLM reasoning via on-policy multilingual thinking language selection during RL, outperforming English-only training and extending exploration with diverse language preferences.
A parameter-efficient plugin extends frozen 10-second ECG foundation models to long, variable-length recordings via compatible long-sequence processing and semantically informed temporal modeling, outperforming sliding-window and pooling baselines.
FRUC enables feedforward, calibration-free dynamic 3D Gaussian reconstruction from uncalibrated multi-vehicle views via an ego-centric occlusion field and residual cross-agent fusion, achieving state-of-the-art rendering quality and efficiency.
Reformulating quality-diversity optimization as multi-objective optimization with many objectives enables set-based scalarization methods to solve QD problems with theoretical guarantees and competitive performance.
ODEWorld learns continuous latent velocity fields via ODEs to enable arbitrary-resolution world modeling, solving representation collapse and excelling at video generation and robotic control.
Frontier LLMs suffer Internal Safety Collapse, generating harmful content during benign tasks with 95.3% failure rates and revealing alignment does not eliminate underlying risks.
VEX-Bench benchmarks verification complexity of LLM-generated misinformation, showing high-VEX false content costs 3-169x less to create than to verify and risks misallocating scarce screening resources.
VisHarness trains a visual agent to orchestrate heterogeneous experts for multi-turn reasoning, achieving strong results on segmentation, detection, and counting tasks.
ROCKET aligns multiple VLA layers to a 3D vision model via residual streams and shared projectors, achieving near-state-of-the-art LIBERO success with about 4% compute.
Delta-Adapter extracts a semantic delta from single image pairs to train exemplar-based editors without paired examples, improving accuracy and generalization.
Platonic Representation Defense detects and purifies backdoored self-supervised encoder representations via cross-model energy functions without labels or training data. It substantially improves robustness across multiple encoders and over ten attacks in fully black-box settings.
A language-assisted clustering framework uses cross-modal relational signals and adaptive semantic centers to improve clustering accuracy by 2.6% over state-of-the-art methods.
Freezing and reusing source latent spaces in diffusion models under distribution shift causes score errors governed by subspace misalignment and amplified ambient noise, with shared representations needed for reliable generation.
FoMEMO proposes foundation models for expensive multi-objective optimization that use synthetic pre-training and in-context preference-conditioned posteriors to optimize unknown problems without further training.
OmniTraffic introduces a controllable 3D traffic generation pipeline and benchmark with 8M VQA samples for spatio-temporal reasoning, revealing large model gaps and improved real-world performance via simulated fine-tuning.