CASE contextualizes tabular embeddings via a dataset-anchored Gemma 3 language model to resolve feature semantics, substantially improving tabular learner accuracy especially with scarce data.
FlexTab uses a shared encoder and task-specific decoders for in-context tabular learning, achieving state-of-the-art results on classification, regression, anomaly detection, and entity matching.
STRABLE introduces 108 real-world string-and-number tables and benchmarks 445 pipelines, finding simple embeddings with advanced learners suffice for categorical tables while LLMs help on free-text tables.
This paper benchmarks 2D tabular attention across GPU backends, finding optimal choices vary by row versus column attention, hardware, and sequence length.