Good Papers

Showing Tokenization Show all papers

45%Niche pick
?Niche pickVote to see the score

Comparing Transformers and Hybrid Models at the Token Level

Yanhong Li, Will Merrill

Atlanta Poster Session 2, Wed, Dec 9, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Model-Aware Tokenizer Transfer

Mykola Haltiuk, Aleksander Smywiński-Pohl

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

SAVE: Sparsity-Aware Influence Estimation for Vocabulary-Expanded LLMs

Seungyoo Lee, Giung Nam, Seanie Lee, Hyungi Lee and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Joint Sequence--Vocabulary Selection for Efficient LLM Distillation

Xueli Geng, Weicheng Zhao, Xutong Mu, Tianrui Wei and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

MoTo: Mixture of Tokenizers Towards Fair Multilingual Language Modeling

Gül Sena Altıntaş, Colin Raffel

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Not All Tokens Should Be Treated Equally: Context Credits Reassignment

Shujian Gao, Yuan Wang, Jiamei Yan, Penghao Zhou and 4 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

The Communication Bottleneck: A Round-Trip Study of Compositional Serialization in Language Models

Xavier Suau, Alex Ferrando, Luca Zappella, Samy Bengio

Paris Poster Session 3, Thu, Dec 10, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

BitMTP: When Multi-Token Prediction Meets Low-Bit Large Language Models

Ning Zhang, Shihao Wang, Jinrui Zhang, Chaodong Xiao and 1 more

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
86%Must read
?Must readVote to see the score

Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities

Universal Byte-Level Encoding routes 3, 4 byte UTF-8 via UTF-16 to lower token counts for high-premium scripts without raising costs for efficient spans, reducing cross-lingual token-budget disparity while preserving model quality.

Hyunsik Kim, Youngmoon Jung

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 0/5
83%Must read
?Must readVote to see the score

Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

Per-token LLM billing is unauditable because providers control the evidence, enabling hidden reasoning inflation up to 1,469% and tokenization-based over-reporting of 50.85% without detection.

Shahinul Hoque, Jinghuai Zhang, Jinyuan Stella Sun, Fnu Suya

Atlanta Poster Session 3, Thu, Dec 10, 10:00 AM–1:00 PM, Hall C1 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5
80%Must read
?Must readVote to see the score

SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation

SimCT recovers lost cross-tokenizer supervision by comparing multi-token continuations in on-policy distillation, improving reasoning and code generation over exact token matching.

Jie Sun, Mao Zheng, Mingyang Song, Qiyong Zhong and 5 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 0/5
83%Must read
?Must readVote to see the score

Effective Context in Transformers: An Analysis of Fragmentation and Tokenization

Smaller units can worsen prediction via fragmentation despite larger windows, while greedy tokenization extends effective context via compression, yielding an information-theoretic framework for representation choices.

Amirmehdi Jafari Fesharaki, Mohammadamin Rami, Aslan Tchamkerten

Paris Poster Session 1, Wed, Dec 9, 12:30 PM–2:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 2/5