CLIOPATRA attacks privacy-preserving LLM insight platforms with malicious chats to leak target medical histories with nearly 100% precision in 65% of cases, showing layered heuristic protections are insufficient.
RefGRPO closes LLM agents' reflection gap via a free calibration bonus and dynamic schedule, improving calibration and task accuracy. Calibrated reflections enable self-improvement without outcome supervision and effective selective prediction.
A unified framework combines unsupervised pretraining and supervised neural knowledge graph learning, with a nonasymptotic risk bound showing unlabeled data reduces downstream prediction error.
Removing hindsight relabeling enables stable pure-RL online finetuning of Decision Transformers via adapted GRPO with sub-trajectory optimization and active sampling, achieving state-of-the-art results.
MT-JailBench provides a modular framework for comparing multi-turn jailbreak attacks under standardized conditions, finding that prompt generation drives success and recomposed components yield stronger attacks.