Good Papers

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

BERT pre-trains deep bidirectional Transformers using masked language modeling and next sentence prediction, achieving state-of-the-art results across language understanding tasks via fine-tuning.

Jacob Devlin, Ming‐Wei Chang, Kenton Lee, Kristina Toutanova

Published 201934,030 citationsPaper ↗

89%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel16/21reviewers recommend it
lenient 4/5
medium 9/11
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
BERT's bidirectional pre-training and missing-recipe reproducibility make it a foundational must-read, though its ablations rarely isolate direction from scale and citations mostly trade on a false aura of completeness.

Abstract

Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 2019.