69%Highly rated
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
RealCompanion benchmarks AI companions on longitudinal real-world chats, finding needed past messages are usually recent, memory detectors fail on real messages, and persona reconstruction costs vary 31-fold at equal F1.
Published Oct 1, 2026 · 0 citations · ▲ 268 on Hugging Face
0% Readers0 of 1 upvoted
13/20 AI panelreviewers recommend it
Only vote on papers you've read. Sign in with GitHub to vote.
AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5