MulTaBench benchmarks 40 multimodal tabular datasets and shows target-aware tuning of text and image embeddings improves predictive performance over frozen embeddings.
Low-rank adaptation regularizes critic learning by constraining updates to low-dimensional subspaces via frozen base weights, reducing loss and improving off-policy RL performance.
INTRA unifies retrieval and generation via decoder attention over internal encoder states, outperforming engineered RAG on question-answering benchmarks.
Normalizing weights and hidden states to the unit hypersphere makes nGPT robust to 4-bit arithmetic, enabling stable end-to-end NVFP4 training without extra scaling or transforms.