ECC calibrates semantic embeddings with limited model comparisons to cluster queries by latent capability demands, improving LLM ranking by ~18 points over semantic baselines and aiding query routing.
AgentOWL jointly learns hierarchical neural options and an abstract world model for sample-efficient skill acquisition, outperforming baselines on object-centric Atari games with fewer samples and stronger generalization.
LLM pipeline evaluation variance is underestimated because design choices are ignored, so corrected intervals restore coverage and cut benchmark gaming.
A tri-modal masked diffusion model pretrained from scratch on text, image-text, and audio-text data achieves strong cross-modal generation and introduces an SDE-based batch-size reparameterization.
Second-order VAW on short ARX models achieves dimension-free regret O(δ⁻⁴ log² T) for marginally stable linear sequence prediction via universal preconditioning and Faber polynomial analysis.
MIND uses sliced Wasserstein distance via sorting to evaluate generative models with 10x better sample efficiency, 100x faster computation, and greater adversarial robustness than FID.