ComPose tracks 6DoF object pose in RGB video by using hand motions as complementary cues, achieving robust accuracy under severe occlusion without external priors.
Feed-forward framework decomposes scenes into instance-structured 3D token groups from unposed multi-view images to enable reconstruction, segmentation, and direct object editing.
NVQLS proposes a hybrid quantum-classical unsupervised operator learning framework using a Legendre-Galerkin formulation to solve parametric PDEs with improved accuracy and theoretical speedups.
RelationVGGT enables feed-forward 3D spatial relation segmentation across multi-view images without camera poses or category names by combining visual semantics with geometry-aware representations via a relation transformer.
An anchor-projected framework maps hidden representations into shared coordinates to extract and transfer behavioral directions across model families without fine-tuning, achieving high cross-family steering and detection accuracy.
RADS applies reachability analysis and constrained reinforcement learning to steer diffusion trajectories away from memorized outputs via caption embedding perturbations, improving diversity, quality, and alignment without altering the model.
Out-of-distribution shifts rotate active subspaces, misaligning dictionary explainers; a geometry-adaptive realignment using unlabeled OOD activations closes the faithfulness gap and restores causal interpretability without training.
Cross-modal sparse autoencoders exhibit feature heterogeneity where shared concepts activate different latents across image and text modalities, and training modality-specific autoencoders with post-hoc alignment improves reconstruction, retrieval, and steering.
Task vector design via distributional alignment with in-context learning minimizes next-token probability discrepancy, yielding a linear method that improves accuracy by 9.2% and enables cross-scale transfer.
ARFP integrates key-bound face cloaking with adversarial restoration-aware training to resist inverse purification attacks while allowing authorized reversible recovery and tamper detection.
TRACE introduces tourism dialogues pairing multi-turn recommendations with review citations and rejection turns to expose the Three-Competency Gap across accuracy, grounding, and recovery.
SAEParate clusters diffusion latent features by concept via contrastive learning to improve precise concept unlearning with reduced cross-concept interference.
MahaVar detects OOD inputs by measuring class-wise Mahalanobis distance variance, which is high for in-distribution samples due to Neural Collapse geometry, achieving state-of-the-art results on CIFAR-100 and ImageNet.
Quantum autoencoders achieve optimal blind single-copy quantum compression with k encoder and n decoder ancillas, pinpointing the universal encoder threshold and showing isometric decoders are nearly optimal.
A new benchmark reveals current multimodal AI fails at multi-step ECG reasoning, achieving near-zero completion in linking clinical criteria to visual signal evidence.
Diffusion models memorize by overestimating training samples during early denoising, collapsing latent trajectories and accelerating convergence to memorized images via classifier-free guidance.
AET framework unifies nonlinear multi-objective RL under SER and ESR via aggregation-expectation-transformation decomposition, and AETDICE enables offline optimization via augmented density-ratio estimation.