Intern-Atlas builds a methodological evolution graph from over one million AI papers to model how methods emerge and adapt, enabling automated idea evaluation and generation.
SAGAS learns a reusable latent reachability graph from fixed offline trajectory fragments to synthesize cost-aware accepting plans for unseen linear temporal logic specifications via test-time semantic graph augmentation and Büchi search without online interaction or retraining.
TACache decomposes rectified flow velocity errors into magnitude and direction components to skip steps and reconstruct velocities without extra evaluations, achieving up to 4.14x faster image and 2.11x faster video generation.
MemForest partitions agent memory into event trees and progressively merges redundant nodes to cut storage and retrieval costs while preserving nearly all performance.
WorldAct converts static generated 3D worlds into editable, interaction-ready scenes via multimodal decomposition and object reconstruction to enable manipulation and embodied tasks.
D²Quant improves sub-4-bit LLM weight-only quantization via dual-scale quantizers for down-projection matrices and deviation-aware LayerNorm correction, boosting accuracy without extra bit budget.
AdaCodec uses predictive visual codes to send full reference frames only when unpredictable, cutting video MLLM tokens by 7x while improving long-video benchmark scores and reducing latency.
DeformGen uses dynamics-based topological augmentation to generate diverse deformable object states and warp trajectories for improved manipulation policy learning.
RePO-VLA improves vision-language-action robustness by assigning roles to success, recovery, and failure trajectories, raising adversarial success from 20% to 75%.
VideoOdyssey benchmarks ultra-long video understanding via continuous certificates averaging 16 minutes, revealing MLLM failures in continuous reasoning and omni-modal perception.
HandEdit provides a 200M-instance benchmark and dataset for transforming egocentric human hands into diverse dexterous robot embodiments via image editing. It evaluates 11 baselines across hand-only and hand-arm tracks with embodiment-aware metrics.
SkeMex improves medical agents via self-evolving skill memory that distills reusable procedural knowledge, governs retention by utility, and outperforms memory-based agents across clinical tasks.
MultiTalk introduces 57.6k hours of synthetic multi-party bilingual dialogue data and MultiTalkBench for long-form full-duplex evaluation, training a model that sustains coherent extended multi-party English-Chinese conversation and outperforms open-source baselines.
IndustryCode is a multi-domain, multi-language benchmark of 579 industrial coding sub-problems; Claude 3.5 Opus reaches 68.1% sub-problem and 42.5% main-problem accuracy.
BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
PO-PDDL learns symbolic POMDPs from robot videos to enable belief-space planning under partial observability and stochasticity, outperforming prior methods with lower planning cost.
BAS-VLA calibrates frozen VLA actions via breaking-centered calibration and selective preservation gating to suppress stale-task drift and separate semantics. It achieves 98% clean success, 0% under target swaps, and 70% under style shifts versus 42%.
BitDance is an autoregressive image generator that predicts binary visual tokens via a diffusion head and next-patch decoding, achieving state-of-the-art FID with far fewer parameters and much faster inference.
DiPO disentangles perplexity into exploration and exploitation subspaces to enable fine-grained trade-offs, improving LLM reasoning and function calling via stable perplexity-guided policy optimization.
xHC expands Transformer hyper-connections beyond four streams via sparse updates and temporal augmentation, improving scaling efficiency. It boosts 18B MoE downstream scores by 4.0 points over mHC with lower compute and reduced memory traffic via xHC-Flash.
Winfree Oscillatory Neural Network applies generalized synchronization dynamics to vision and reasoning tasks, scaling to ImageNet-1K and achieving 80.1% Maze-hard accuracy with 1% of prior parameters.
ForceFlow uses force-aware flow matching with asymmetric multimodal fusion and vision-to-force handover to achieve robust contact-rich manipulation with 37% higher success and stronger zero-shot generalization.
Flash-KMeans eliminates GPU HBM bottlenecks via fused assignment and inverse mapping updates, delivering up to 17.9x speedups over existing exact k-means implementations.
FocusDepth uses spatially-aligned multi-scale prompt fusion to boost target-region depth accuracy and sharp boundaries while preserving global geometry, outperforming global baselines on FDE-Bench.
HOMIE unifies inter- and intra-subject video personalization via multimodal guidance and reference embeddings, achieving state-of-the-art human-object interaction fidelity.
ViDiHand leverages pretrained video diffusion models to reconstruct 4D hand poses directly from full egocentric video without detectors, substantially outperforming prior methods on ARCTIC, HOT3D, and HOI4D.
CoPhy distills vision-language cognition into a BEV encoder and pairs it with an auto-regressive world model for action-conditioned forecasting to enable reinforcement learning with dual physical and cognitive rewards, achieving state-of-the-art autonomous driving results.
AdaViG uses internal generation-intent and visual-fidelity signals to abort unhelpful visual reasoning steps early, improving multimodal reasoning accuracy by up to 5.7% while cutting visual generation costs by 25-91%.
Astra enhances vision-language model spatial reasoning by letting agents generate imagined simulator views via RL, improving MMSI-Bench scores over direct answering.
AMS replaces global token eviction with adaptive region-aware KV quotas to prevent reasoning block wipe-out, boosting long-context performance without extra attention overhead.
Generation Navigator is a state-aware multi-turn text-to-image agent that learns to steer generation via trajectory-level reinforcement learning, achieving a 0.90 WISE score and 79.06% reasoning accuracy.
A diagnostic taxonomy maps AVLM failure signatures to targeted development interventions, enabling traceable industry-scale video moderation system improvements.
Visual-ERM is a multimodal generative reward model that evaluates vision-to-code outputs in rendered visual space, improving Qwen3-VL-8B-Instruct by up to 8.4 points and outperforming larger models on fine-grained visual discrepancy benchmarks.
Justitia schedules task-parallel LLM agents via memory-centric cost prediction and virtual-time fair queuing to improve efficiency while preserving fairness and worst-case delays.
V-CAST prunes video tokens via curvature-guided temporal budgets and dual-anchor spatial selection, achieving 98.6% original performance with 86.4% latency.
EgoTac predicts tactile signals from egocentric videos using 5.7M image-tactile pairs, achieving under 0.06N force error and outperforming contact estimators.
A teacher-aware evolutionary framework uses learned optimization policies as behavioral teachers to evolve static executable heuristics, improving combinatorial optimization benchmarks without neural inference at deployment.
ViCO minimizes vision tokens via consistency training across compression ratios, cutting tokens up to 50% while preserving capabilities through semantic-based routing.
GMOS grounds moving object segmentation in 3D space and time using an RGB video framework, achieving state-of-the-art results across MOS benchmarks with faster online inference.
CoWorld-VLA embeds multi-expert world tokens into vision-language-action models and couples diffusion planning with scene context to generate continuous ego trajectories, improving autonomous driving performance.
G²TR uses generation-branch signals to reduce visual tokens in unified multimodal models, cutting prefill computation by 1.94× while preserving reasoning and editing performance.
EasyLens is a training-free plug-and-play module that amplifies subtle lesion representations in frozen medical vision-language models via prototype-based patch selection and morphology-guided residual enhancement, improving detection across datasets.