StepSTEM introduces 283 graduate-level STEM problems with strictly complementary visual-textual inputs to evaluate cross-modal reasoning via step-level alignment, revealing current MLLMs achieve only 38.29% accuracy due to heavy reliance on text.
MIRAGE learns continuous latent reasoning for mobile agents, cutting decoded tokens 75% while matching explicit chain-of-thought accuracy and improving baselines up to 10.2 points via generative world modeling.