ComPose tracks 6DoF object pose in RGB video by using hand motions as complementary cues, achieving robust accuracy under severe occlusion without external priors.
Diagonal regularizers saturate at a prior-driven power law because isotropic truncation noise and eigenvalue counting yield flat loss landscapes, so learned diagonal forms barely beat the closed form and only cross-mode coupling enables real gains.
A reasoning-prefix masking framework distills think-answer visual reasoning into compact VLMs by masking salient reasoning cues to force visual anchoring, improving multimodal benchmarks over prior distillation methods.
Random soft prompt injection boosts LLM math reasoning by flattening early token distributions to diversify reasoning paths and widen Pass@N without any training.