POME applies truncated SVD to weight-update differences to equalize dominant directions and prune noise, boosting fine-tuned LLM performance by up to 2.5% with no extra cost.
FlexMoE ranks and prunes Mixture-of-Experts channels via discrete actions to generate nested subnetworks across budgets, preserving ~99.8% performance at 50% pruning without fine-tuning.