Difference-informed pruning preserves output differences via difference-aware weight scoring, improving LLM sparsity over activation and reconstruction baselines at minimal cost.
A systematic comparison of general agent architectures finds backbone choice dominates performance while architecture shifts results up to 12pp, and open models suffer generality sinks.
Federated concept-based models aggregate distributed concept annotations across institutions, adapt architectures to evolving supervision, and enable interpretable inference for locally unavailable concepts while preserving privacy.
LLaDA-Guard uses masked diffusion to score responses under each safety label and classify by difference, improving calibration, reducing over-defense, and enabling token-level risk localization with 60.7% prompt rewriting success.
Vanilla LoRA matches variant performance within 1-2% when learning rates are tuned, and differing optimal rates stem from Hessian eigenvalue variations.
DiagnosticIQ benchmarks LLM recommendation of industrial maintenance actions from symbolic rules across 6,690 questions, finding frontier models match human experts but break under structural perturbation due to calibration failures rather than capability gaps.
A generic patch Transformer achieves state-of-the-art zero-shot time series forecasting via simple training, with scaling and data ablations isolating key performance drivers.
HSCO-Bench evaluates LLM agents on end-to-end hardware-software co-design for SoCs, finding only two frontier models generate valid prototypes with suboptimal resource use.