FairMT introduces a unified fairness framework for heterogeneous multi-task learning with partial labels, using asymmetric constraint aggregation and head-aware optimization to improve fairness without sacrificing utility.
xMemory decouples agent memories into reusable components before aggregating them hierarchically, improving retrieval quality and token efficiency over flat RAG.
CAREBench evaluates LLM emotion understanding via appraisal reasoning chains, finding stronger models surpass humans on some tasks but lack reasoning and positive emotion recognition.
ComprExIT fixes structural bottlenecks in LLM context compression via explicit cross-layer feature selection and coordinated transport, improving F1 up to 18.5% with minimal parameters and 2x faster compression.
Test-time personalization samples candidates and selects via reward models, proving logarithmic utility scaling but diagnosing user collapse and query hacking, fixed by probabilistic rewards.
Anchored Bipolicy Self-Play uses frozen-base LoRA adapters to separate attacker and defender roles, preventing self-consistency collapse and improving safety with 100x greater parameter efficiency.
Poisoning LLM pretraining requires only ~250 malicious documents regardless of dataset or model scale, revealing constant-cost backdoor injection risks for large models.