Universal Byte-Level Encoding routes 3, 4 byte UTF-8 via UTF-16 to lower token counts for high-premium scripts without raising costs for efficient spans, reducing cross-lingual token-budget disparity while preserving model quality.
Per-token LLM billing is unauditable because providers control the evidence, enabling hidden reasoning inflation up to 1,469% and tokenization-based over-reporting of 50.85% without detection.
SimCT recovers lost cross-tokenizer supervision by comparing multi-token continuations in on-policy distillation, improving reasoning and code generation over exact token matching.
Smaller units can worsen prediction via fragmentation despite larger windows, while greedy tokenization extends effective context via compression, yielding an information-theoretic framework for representation choices.