Replacing dense orthogonal matrices with sign-randomized DCTs accelerates Kashin-based LLM quantization to O(N log N) with guaranteed convergence, achieving 4-bit accuracy competitive with OPTQ and QuIP while maintaining numerical stability and native 2-bit hardware compatibility.
Self-supervised pre-training on one real table yields strong tabular transfer, where feature count predicts usefulness and in-context generalization is retrieval-based.
TwinRouterBench introduces static and live dynamic tracks to benchmark LLM routing at agent step-level using deterministic scoring and live execution on SWE-bench.
P2P uses adaptive prompting and L1-regularized regression to build compact LLM ensembles that emulate human preferences at low cost without fine-tuning. It achieves 0.014 test MSE on American Trends Panel surveys for about $0.80 each and outperforms supervised baselines with under 3% of their traini
Medmarks introduces 30 open-source medical benchmarks evaluating 61 LLMs, finding frontier reasoning models lead, proprietary models are more token-efficient, medical fine-tuning helps, and smaller models show answer-order bias.
AmaraSpatial-10K is a 10,000 synthetic 3D asset dataset optimized for deployment, achieving 3.4x CLIP recall over Objaverse and 99.1% physics stability.
Machine learning papers are hard to reproduce due to missing code, so researchers should prioritize verifiable results through concrete checkability improvements.
Architectural topology modifications eliminate Transformer's grokking phase by bounding representations and fixing attention, but only when aligned with task symmetries.
PRE-ACT models accident risk as a continuously evolving signal that increases approaching crashes, enforcing temporal ordering and distance awareness to suppress false alarms and improve anticipation performance.
Language models prefer correct answers because errors are less compressible, not because they detect truth directly, with coherent false rules eliminating accuracy until competing rules restore it.