Parallel-Synthesis lets LLM synthesizers consume parallel agents' KV caches directly via a cache mapper and adapter, matching text synthesis on seven of nine benchmarks while cutting time-to-first-token by 2.5x-11x.
RigidFormer is a transformer that learns mesh-free rigid-body dynamics via object-level anchors and differentiable Kabsch projection, outperforming mesh-based baselines with faster inference and scalability to 200+ objects.
DARLING uses a learned partition function to jointly optimize language model response quality and semantic diversity via reinforcement learning, improving both quality and novelty across creative and math benchmarks.
Medical imaging pretraining reveals asymmetric cross-domain scaling and power-law transfer, yielding optimized data allocations with a hub-and-island structure that improves transfer over proportional sampling by up to 58%.
ECHO-2 is a distributed RL framework that overlaps rollout generation, dissemination, and training with bounded policy staleness to improve cost efficiency while preserving rewards.
A two-stage adapter embeds foundation model predictions into a constrained multinomial logit, guaranteeing cost monotonicity and valid value-of-time estimates while improving choice accuracy by up to 12.8 percentage points.
TextSeal is a localized LLM watermark using dual-key generation and entropy-weighted scoring for robust provenance and distillation detection without inference overhead.
HiLight trains a lightweight actor via reinforcement learning to insert highlight tags around pivotal evidence spans in frozen LLM contexts, boosting reasoning without altering inputs or requiring evidence labels.
A 1MB replay script outperforms frontier agents on static benchmarks because of flawed environment design and evaluation; the paper proposes PRISM principles, DigiWorld, and hierarchical bootstrap aggregation to fix both.
MAGE uses block-diffusion's aligned all-[MASK] queries to select reusable sparse KV subsets, achieving near-lossless accuracy with up to 6.82x speedup at 128K context.
Schema-derived ODS constraints enable small LMs to match or exceed 15B, 34B models on structural MLIR dialects at 8, 25× speed without retraining, though attribute-heavy dialects remain challenging.
RANSAC scoring analytically marginalizes inlier scale via a conjugate prior, yielding a parameter-free score that outperforms threshold-based methods across data regimes with O(N log N) computation.
Residual Quantization maps contexts to discrete additive codes enabling nonlinear contextual bandits with strictly bounded memory, beating linear variants on 11 of 13 datasets and matching heavy retrained baselines with up to 1000x less memory.
TRIBE v2 synthetic fMRI augmentation improves brain-to-image decoding by up to 68%, though optimal synthetic-to-real ratios vary by dataset, and synthetic-only training achieves above-chance zero-shot decoding.
PISCO enables precise video instance insertion via sparse keyframe control while preserving dynamics, achieving monotonic gains with added signals and outperforming editing baselines.
MemReward propagates rewards through a heterogeneous rollout graph to enable LLM reinforcement learning using only 20% ground-truth labels and achieves over 96% of oracle performance.
Watermarking lacks enforceable standards and audit infrastructure, so current implementations serve as symbolic compliance rather than effective AI oversight.
SWE-Protégé trains small language models to selectively seek expert guidance and avoid looping, achieving 42.4% Pass@1 on SWE-bench Verified with minimal expert use.
MetaCanvas enables multimodal LLMs to plan directly in diffusion latent spaces, outperforming global-conditioning baselines across six precise visual generation tasks.
TerminalWorld automatically builds terminal benchmarks from wild recordings, yielding 1,530 tasks where top agents achieve only 62.5% success with weak correlation to expert benchmarks.