A multi-reward RLIF framework combining cluster-voting and self-certainty rewards with GDPO normalization and KL-Cov regularization prevents collapse and matches supervised RLVR performance.
BSTabDiff partitions high-dimensional tabular features into latent blocks with shared subunits to learn global dependencies in compact space, enabling stable synthesis of realistic HDLSS data with copula-driven dependence and explicit missingness.