LEAD adaptively calibrates reasoning length via online self-adaptive rewards, achieving highest accuracy and efficiency scores with shorter outputs than base reasoning models.
Medmarks introduces 30 open-source medical benchmarks evaluating 61 LLMs, finding frontier reasoning models lead, proprietary models are more token-efficient, medical fine-tuning helps, and smaller models show answer-order bias.
PromptMIA uses adversarial soft prompts to exploit federated prompt-tuning updates for highly effective membership inference attacks that bypass standard defenses.
A systematic benchmark shows out-of-distribution detector competitiveness depends mainly on learned representations rather than score design, with neural collapse metrics predicting top detector choices without extra out-of-distribution data.
A retinal model with a novel opsin layer simulates color vision evolution and optimizes task-specific camera spectral filters via mutation-driven adaptation.