AgentBrew learns tool-use policies offline from raw real-world trajectories via retrospective task inference and PMI-based credit assignment, improving Qwen3-32B by +8.7 accuracy over larger baselines.
A unified framework reveals GNN under-confidence stems from final-layer weight decay and node distance, fixed by reducing decay and node-level calibration.