Good Papers

From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models

Pretraining data political biases propagate through language models into unfair hate speech and misinformation detectors, reinforcing polarization.

Shangbin Feng, Chan Young Park, Yuhan Liu, Yulia Tsvetkov

Published 2023138 citationsPaper ↗

74%
OverallHighly rated
?
OverallHighly ratedVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel9/21reviewers recommend it
lenient 5/5
medium 4/11
strict 0/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Valued for tracing pretraining political bias into downstream hate speech and misinformation fairness, the study is undermined by correlational methods that leave label bias unexamined and offer no mitigation or code.

Abstract

Language models (LMs) are pretrained on diverse data sources, including news, discussion forums, books, and online encyclopedias.A significant portion of this data includes opinions and perspectives which, on one hand, celebrate democracy and diversity of ideas, and on the other hand are inherently socially biased.Our work develops new methods to (1) measure political biases in LMs trained on such corpora, along social and economic axes, and (2) measure the fairness of downstream NLP models trained on top of politically biased LMs.We focus on hate speech and misinformation detection, aiming to empirically quantify the effects of political (social, economic) biases in pretraining data on the fairness of high-stakes social-oriented tasks.Our findings reveal that pretrained LMs do have political leanings that reinforce the polarization present in pretraining corpora, propagating social biases into hate speech predictions and misinformation detectors.We discuss the implications of our findings for NLP research and propose future directions to mitigate unfairness. 1 Warning: This paper contains examples of hate speech.