LADA trains vision-language-action models using under 5% language annotations via latent action codebooks, achieving 87.98 driving scores on Bench2Drive.
CurveBench introduces a 756-image benchmark for hierarchical containment reasoning over nested Jordan curves, showing top models achieve only 19% accuracy on hard cases.
RelAgent is an LLM agent that builds SQL feature queries and selects predictive models for relational learning, yielding fast, interpretable predictions deployable via standard databases.
Anchored Bipolicy Self-Play uses frozen-base LoRA adapters to separate attacker and defender roles, preventing self-consistency collapse and improving safety with 100x greater parameter efficiency.