3D-PLOT-LLM inserts learnable part tokens into frozen point features to enable part-level reasoning in 3D LLMs with under 1M parameters, outperforming prior part-aware models on part-QA and grounded description benchmarks.
RLBD trains a reinforcement learning policy to adaptively select Benders cuts, substantially improving computational efficiency and generalizing across problem variations.
SkillsBench benchmarks agent skills across 87 tasks, finding curated skills boost pass rates by 16.6 points, with focused small bundles often outperforming larger ones.
WildBox provides aerial monocular 3D wildlife annotations and benchmarks showing zero-shot 3D detection collapses to zero, with fine-tuning reaching 13.17 AP3D and depth as the dominant failure mode.
Adam converges with high probability on generalized-smooth objectives under only second-moment stochastic gradients, matching a sharp δ^{-1/2} confidence dependence and yielding expectation rates for p<1.
PhenoAIR reformulates Cell Painting mechanism prediction as calibrated evidence reasoning via multi-agent evaluation of noisy retrieved neighbors, outperforming matching and LLM baselines across open-world settings.
This paper introduces a multimodal benchmark for cause-of-death inference in child mortality data, showing zero-shot language models synthesize unstructured medical evidence differently than supervised baselines.
QUEST trains open deep research agents via synthetic rubric-tree tasks and context management, achieving frontier-level performance across eight benchmarks with only 8K examples.
ACuRL enables autonomous continual learning for computer-use agents via curriculum reinforcement learning, yielding 3-29% gains without catastrophic forgetting or human data.