StepSTEM introduces 283 graduate-level STEM problems with strictly complementary visual-textual inputs to evaluate cross-modal reasoning via step-level alignment, revealing current MLLMs achieve only 38.29% accuracy due to heavy reliance on text.
XDecomposer learns prior-free multiphase X-ray diffraction decomposition as set prediction to identify constituent phases and proportions without candidate lists. It improves reconstruction accuracy and phase identification across simulated and experimental datasets.