DualKV eliminates shared-prompt replication in RL training via FlashAttention kernels that process shared and per-sequence KV regions separately, achieving up to 3.82x policy-update speedup.
TabPrep is a lightweight feature engineering pipeline that targets structural data patterns to consistently boost tabular model performance across benchmarks.
A neuro-symbolic framework trains a 4B-parameter model via MCTS-curated preference data and two-stage post-training to learn SAT cubing heuristics matching top symbolic methods.