SCOPE co-evolves a task-generating challenger and retrieval solver with rubric-based self-judging to improve open-ended and QA performance without curated data.
B-XAIC benchmark evaluates explainable AI for graph neural networks on real molecular tasks with ground-truth rationales, revealing major limitations in current explanation methods.