PRPO incorporates column-permutation invariance into LLM post-training via label-preserving permutations and two-level advantage estimation, enabling an 8B model to match specialized tabular baselines and outperform 685B reasoning LLMs by up to 53%.
ROME uses role-playing LLMs to generate questionnaire answers from posts, then routes them via mixture-of-experts to improve personality detection and mitigate label scarcity.
Reformulating neural operators in d+1 dimensions via auxiliary embedding evolution achieves lowest relative L2 error across benchmarks without brute-force scaling.