Good Papers

Lingtai: What Concept Geometry Reveals—and Does Not Reveal—About LLM Inference

Lingtai introduces a training-free concept telemetry layer that reveals inference-time uncertainty-linked activity and execution-specific trajectory structures in LLMs without tracking correctness.

Jiangang Chen

Published Oct 1, 2026arXiv ↗

80%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel12/20reviewers recommend it
lenient 3/5
medium 7/10
strict 2/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
Lingtai delivers a training-free, low-overhead telemetry signal whose uncertainty-linked concept geometry and execution-specific trajectory identity are compelling, yet it explicitly refuses to supply a stable internal correctness coordinate and still needs broader benchmarks and dense-baseline validation.

Abstract

Observing what a large language model computes during autoregressive inference—online and without training probes—remains difficult. We introduce Lingtai, a training-free concept telemetry layer: at each generation step, residual states are projected onto a domain-specific bank of named concept anchors, constructed without labeled concept examples, outcome labels, gradient fitting, or activation-space optimization, producing a structured per-step concept-coordinate signal. Across code generation and grade-school mathematical reasoning, this signal exhibits a robust association with predictive uncertainty: the association survives problem-identity and token-position controls and is not attributable to a single token type, is not explained by a simple correct/incorrect mixture on GSM8K, and is not reproduced by matched random anchors; it is markedly weaker or direction-inconsistent in K-means and PCA projections. Two structures emerge: a recurring uncertainty-linked activity signal whose functional geometry is task-conditioned (distinct activity–entropy shapes on HumanEval, MBPP, and GSM8K), and an execution-specific trajectory identity with strong local inertia but weak re-instantiation invariance—under completion-only elastic alignment, corruption at k=32 (approximately a median quarter of the completion) on the matched re-execution subset still retrieves the archived episode at 62.0%, while a fresh execution retrieves it only 11.7–16.0% of the time. Finally, a matched audit finds no evidence that the scalar concept-activity signal used here supplies a stable correctness coordinate under the tested protocol; we therefore treat correctness as externally supplied. Telemetry adds 0.7–1.6% per-token decode overhead for the 161-anchor code implementation, with unchanged generated tokens.