Good Papers

What Gets Measured Gets Managed: Sign-aware Recommendation Needs Sign-aware Evaluation

Sign-aware recommender systems embed valence but rank blindly, hidden by metrics that ignore disliked items; proposed signed metrics expose poor valence protection and provide trainable fixes.

Minchan Kim, Jungmin Hwang, Hyunwoo Park

Published 2026Sydney Poster Session 4 · Wed, Dec 9, 5:00 PM–8:00 PM local time · Hall 1-4arXiv ↗OpenReview ↗

86%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel14/20reviewers recommend it
lenient 5/5
medium 7/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Sign-aware recommender systems have recently been developed to leverage negative feedback for a deeper understanding of user preferences. However, our empirical diagnosis reveals that state-of-the-art graph-based sign-aware recommender systems are paradoxically valence-blind. Even though they explicitly incorporate sign information during training, they consistently fail to differentiate liked items from disliked ones at the ranking stage, frequently infiltrating top-K recommendations with disliked content. Through linear probing, we show that while valence information exists in the learned embeddings, it remains inaccessible to the inner-product scoring function. This widespread failure remains entirely undetected because conventional evaluation metrics, such as Recall, HR, and NDCG, assign a uniform utility of zero to both negative and unobserved items, creating a systematic evaluation blind spot. To bridge this gap, we propose a family of signed metrics, Signed Recall, Signed HR, and Signed NDCG, that explicitly penalize the recommendation of disliked content. Systematic re-evaluation under our proposed metrics fundamentally reshapes the established performance landscape, revealing that methods ranked highly under conventional metrics often fail to protect users from disliked content. Finally, through a proof-of-concept auxiliary loss, we confirm that the proposed metrics provide actionable training signals, guiding models toward valence-aware behavior without sacrificing conventional relevance. For transparency, our source code is available at: https://anonymous.4open.science/r/signed-rec-benchmark-07E4