PAGER closes the semantic-execution gap for point-precise geometric GUI control via dependency-structured planning and pixel-level execution, achieving 4.1x higher task success than general baselines.
PhotoFlow uses a Director-Reviewer-Reflector agent for closed-loop camera search to generate language-conditioned virtual photographs in arbitrary 3D scenes, outperforming baselines on quality, alignment, and success rate.
Defense-as-Skill implements runtime guard SkillSonar as an editable skill that checks actions against task boundaries, reducing attack success rates substantially across agents via evolved guard-skill optimization.
SciHazard benchmarks LLM scientific safety risks via decomposed harm scoring across 3,000 real-world grounded queries, finding deep research agents 32.3% more harmful than standard models.
RankE co-evolves discrete text-to-image policy and decoder via alternating optimization to eliminate latent covariate shift, improving both FID and CLIP scores.
Visual-ERM is a multimodal generative reward model that evaluates vision-to-code outputs in rendered visual space, improving Qwen3-VL-8B-Instruct by up to 8.4 points and outperforming larger models on fine-grained visual discrepancy benchmarks.
SCHEMA evaluates scientific agent hallucinations via topology-aware diagnostics, showing errors cluster at connected knowledge hubs and correct answers often stem from flawed reasoning.