Good Papers

Probing for Long-Horizon Deductive Reasoning Capabilities in Language Models with Prolog

ProloNg tests long-horizon Prolog reasoning in LLMs, finding most models drop to near-chance accuracy beyond reasoning depth 10.

Hadeel Al-Negheimish, Jasna Ilieva, Yoon Kim

Published Oct 8, 2026arXiv ↗

75%
OverallHighly rated
?
OverallHighly ratedVote to see the score
Readers
–

Only vote on papers you've read. Sign in to vote.

AI panel10/20reviewers recommend it
lenient 4/5
medium 3/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Current frontier LLMs can theoretically process long contexts with 1M tokens or more. But to what extent can they go beyond simple retrieval and perform deeper reasoning over such long contexts? We empirically investigate long-horizon reasoning capabilities of LLMs, focusing on deductive logic expressed in Prolog. We construct ProloNg, a synthetic testbed to probe Prolog Long Reasoning, which systematically varies the complexity (reasoning depth) of problems, where the hardest case has a reasoning depth of 22 and 62k context length. We study 8 reasoning models across 5 families of frontier LLMs, and find that performance degrades substantially as reasoning depth grows, with the majority of models approaching chance beyond depth 10.