Intern-Atlas builds a methodological evolution graph from over one million AI papers to model how methods emerge and adapt, enabling automated idea evaluation and generation.
SAGAS learns a reusable latent reachability graph from fixed offline trajectory fragments to synthesize cost-aware accepting plans for unseen linear temporal logic specifications via test-time semantic graph augmentation and Büchi search without online interaction or retraining.
TACache decomposes rectified flow velocity errors into magnitude and direction components to skip steps and reconstruct velocities without extra evaluations, achieving up to 4.14x faster image and 2.11x faster video generation.
MemForest partitions agent memory into event trees and progressively merges redundant nodes to cut storage and retrieval costs while preserving nearly all performance.
WorldAct converts static generated 3D worlds into editable, interaction-ready scenes via multimodal decomposition and object reconstruction to enable manipulation and embodied tasks.
D²Quant improves sub-4-bit LLM weight-only quantization via dual-scale quantizers for down-projection matrices and deviation-aware LayerNorm correction, boosting accuracy without extra bit budget.
AdaCodec uses predictive visual codes to send full reference frames only when unpredictable, cutting video MLLM tokens by 7x while improving long-video benchmark scores and reducing latency.
DeformGen uses dynamics-based topological augmentation to generate diverse deformable object states and warp trajectories for improved manipulation policy learning.
RePO-VLA improves vision-language-action robustness by assigning roles to success, recovery, and failure trajectories, raising adversarial success from 20% to 75%.
VideoOdyssey benchmarks ultra-long video understanding via continuous certificates averaging 16 minutes, revealing MLLM failures in continuous reasoning and omni-modal perception.
HandEdit provides a 200M-instance benchmark and dataset for transforming egocentric human hands into diverse dexterous robot embodiments via image editing. It evaluates 11 baselines across hand-only and hand-arm tracks with embodiment-aware metrics.
SkeMex improves medical agents via self-evolving skill memory that distills reusable procedural knowledge, governs retention by utility, and outperforms memory-based agents across clinical tasks.
MultiTalk introduces 57.6k hours of synthetic multi-party bilingual dialogue data and MultiTalkBench for long-form full-duplex evaluation, training a model that sustains coherent extended multi-party English-Chinese conversation and outperforms open-source baselines.
IndustryCode is a multi-domain, multi-language benchmark of 579 industrial coding sub-problems; Claude 3.5 Opus reaches 68.1% sub-problem and 42.5% main-problem accuracy.
BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
PO-PDDL learns symbolic POMDPs from robot videos to enable belief-space planning under partial observability and stochasticity, outperforming prior methods with lower planning cost.
BAS-VLA calibrates frozen VLA actions via breaking-centered calibration and selective preservation gating to suppress stale-task drift and separate semantics. It achieves 98% clean success, 0% under target swaps, and 70% under style shifts versus 42%.
BitDance is an autoregressive image generator that predicts binary visual tokens via a diffusion head and next-patch decoding, achieving state-of-the-art FID with far fewer parameters and much faster inference.
DiPO disentangles perplexity into exploration and exploitation subspaces to enable fine-grained trade-offs, improving LLM reasoning and function calling via stable perplexity-guided policy optimization.
xHC expands Transformer hyper-connections beyond four streams via sparse updates and temporal augmentation, improving scaling efficiency. It boosts 18B MoE downstream scores by 4.0 points over mHC with lower compute and reduced memory traffic via xHC-Flash.
Winfree Oscillatory Neural Network applies generalized synchronization dynamics to vision and reasoning tasks, scaling to ImageNet-1K and achieving 80.1% Maze-hard accuracy with 1% of prior parameters.
ForceFlow uses force-aware flow matching with asymmetric multimodal fusion and vision-to-force handover to achieve robust contact-rich manipulation with 37% higher success and stronger zero-shot generalization.
Flash-KMeans eliminates GPU HBM bottlenecks via fused assignment and inverse mapping updates, delivering up to 17.9x speedups over existing exact k-means implementations.
FocusDepth uses spatially-aligned multi-scale prompt fusion to boost target-region depth accuracy and sharp boundaries while preserving global geometry, outperforming global baselines on FDE-Bench.
HOMIE unifies inter- and intra-subject video personalization via multimodal guidance and reference embeddings, achieving state-of-the-art human-object interaction fidelity.
ViDiHand leverages pretrained video diffusion models to reconstruct 4D hand poses directly from full egocentric video without detectors, substantially outperforming prior methods on ARCTIC, HOT3D, and HOI4D.