PhenoAIR reformulates Cell Painting mechanism prediction as calibrated evidence reasoning via multi-agent evaluation of noisy retrieved neighbors, outperforming matching and LLM baselines across open-world settings.
QUEST trains open deep research agents via synthetic rubric-tree tasks and context management, achieving frontier-level performance across eight benchmarks with only 8K examples.
ACuRL enables autonomous continual learning for computer-use agents via curriculum reinforcement learning, yielding 3-29% gains without catastrophic forgetting or human data.