MulTaBench benchmarks 40 multimodal tabular datasets and shows target-aware tuning of text and image embeddings improves predictive performance over frozen embeddings.
Under structured cross-modal nuisance correlation, cross-modal alignment and prediction have complementary failure modes partitioned by separation ratios into four regimes, with a data-driven procedure identifying preferred objectives and when neither beats single-modality baselines.
Applying CVaR to the immediate belief cost targets per-step state uncertainty while preserving standard MDP structure, enabling any expectation-based planner to become risk-sensitive with unchanged algorithms and end-to-end finite-time guarantees.
KV-compressibility is a learnable property, so KV-CAT trains transformers via masked KV slots to yield representations more amenable to post-hoc compression without sacrificing quality.
SP-CACW minimizes an upper bound on a target client's convergence error via convergence-aware weighting that trades peer bias against variance and excludes harmful peers.
VideoMDM trains 3D human motion diffusion models solely from 2D video poses via depth-weighted reprojection, nearly matching fully 3D-supervised quality without ground-truth 3D data.
For every k, k-WL fails to distinguish some non-isomorphic simple-spectrum graphs, so PRiSM provides the first complete canonicalization of their eigendecompositions to enable universal approximation.
A framework combines multi-armed bandits with low-rank predictions to build doubly robust estimators and valid confidence intervals for identifying the best LLM with fewer evaluations.
A single sensory-prediction recurrent network with Dale's Law co-emerges grid and place cells without supervision, reproducing key spatial coding phenomena.