A geometric framework defines concept frustration as contradictions from missing concepts and detects it in foundation model embeddings to align human and machine reasoning.
MyoChallenge 2025 benchmarks musculoskeletal sports control via simulated table tennis and soccer tasks, advancing agile motor algorithms across 70 teams.
FairMT introduces a unified fairness framework for heterogeneous multi-task learning with partial labels, using asymmetric constraint aggregation and head-aware optimization to improve fairness without sacrificing utility.
P2T uses reference patches as privileged supervision to curate shorter, grounded agent trajectories via bi-objective optimization, improving SWE-bench Pass@1 by up to 10.8 points with ~15% lower inference cost.
A mean-field framework formulates inference-time diffusion control via weighted interacting particles to target distribution-level rewards with theoretical guarantees.
xMemory decouples agent memories into reusable components before aggregating them hierarchically, improving retrieval quality and token efficiency over flat RAG.
CAREBench evaluates LLM emotion understanding via appraisal reasoning chains, finding stronger models surpass humans on some tasks but lack reasoning and positive emotion recognition.
ComprExIT fixes structural bottlenecks in LLM context compression via explicit cross-layer feature selection and coordinated transport, improving F1 up to 18.5% with minimal parameters and 2x faster compression.
Anchored Bipolicy Self-Play uses frozen-base LoRA adapters to separate attacker and defender roles, preventing self-consistency collapse and improving safety with 100x greater parameter efficiency.