Good Papers

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Apodex 1.1 scales agentic intelligence via environment and coordination scaling to achieve leading complex-work performance with smaller models.

B. An, B. An, B. Wang, B. L. Wang, B. Zhang, Feng C, C. Wei, C. Xue, C. Zhang, C. Zhang, D. Ng, E. Min, F. Chen, F. Liu, F. Yang, F. Ye, G. Sun, H. Xu, H. Xu, H. Yang, H. Ye, H. Zhao, H. Zhao, J. Lin, J. Xia, J. Xia, K. Jin, Wang, K., K. Yang, L. Bing, L. Lei, L. Su, Lu Wang, N. Wang, N. Wang, Q. Ren, Q. Yang, R. Li, S. Bai, S. Du

Published Aug 24, 2026▲ 212 on Hugging FaceCode ★ 5,146arXiv ↗

72%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel8/20reviewers recommend it
lenient 5/5
medium 3/10
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Apodex 1.1 earns praise for grounding agentic intelligence in sustained, verifiable working capability with local 35B deployment and AgentOS provenance, yet remains hampered by vague replanning, missing failure-rate specifics, and unnamed baselines.

Abstract

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of executable file, search, and code environments, while \emph{Agentic Coordination Scaling} trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form. These results ground agentic intelligence in useful, verifiable work completed over time and advance our goal of building a \emph{Heavy-Duty Solver} for ambitious, long-running tasks.