BALTO applies balanced token-level credit assignment to mitigate LLM hallucinations by redistributing probability from unsupported to faithful content, outperforming response-level methods on faithfulness benchmarks.
ECHO-2 is a distributed RL framework that overlaps rollout generation, dissemination, and training with bounded policy staleness to improve cost efficiency while preserving rewards.
TrajWiki represents memory as source-grounded evolution trajectories with claim-level updates and a wiki layer to improve long-horizon dialogue performance and interpretability.
Policy mirror descent with inexact TD actor-critic converges for entropy-regularized MDPs in general spaces, with sublinear or linear rates under sufficient TD steps.