In Bayesian multi-armed bandits, pretraining on expert data tightens regret bounds by mutual information with the optimal action, while an information-directed rule selects data sources maximizing immediate information gain, and trust inference safeguards against ineffective or compromised experts.
Anchored Bipolicy Self-Play uses frozen-base LoRA adapters to separate attacker and defender roles, preventing self-consistency collapse and improving safety with 100x greater parameter efficiency.