Kimi K2.5: Visual Agentic Intelligence
Kimi K2.5 is an open-source multimodal agentic model using joint text-vision optimization and Agent Swarm to achieve state-of-the-art agentic, coding, vision, and reasoning results with up to 4.5x lower latency.
Published Feb 2, 20262 citations▲ 277 on Hugging FaceCode ★ 2,313arXiv ↗

Only vote on papers you've read. Sign in with GitHub to vote.
Kimi K2.5 delivers a compelling agent swarm with verifiable latency gains and a useful checkpoint release, though its joint vision-language optimization remains vague and its agentic gains are hard to separate from orchestration.
Abstract
We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5 introduces Agent Swarm, a self-directed parallel agent orchestration framework that dynamically decomposes complex tasks into heterogeneous sub-problems and executes them concurrently. Extensive evaluations show that Kimi K2.5 achieves state-of-the-art results across various domains including coding, vision, reasoning, and agentic tasks. Agent Swarm also reduces latency by up to $4.5\times$ over single-agent baselines. We release the post-trained Kimi K2.5 model checkpoint to facilitate future research and real-world applications of agentic intelligence.