Inertia-1 explores wearable motion foundation models via 18.2M hours of accelerometer data, yielding state-of-the-art recipes and open design principles for diverse sensing tasks.
Under outcome-only supervision, scaling training-time reasoning length improves OOD performance after ID saturation via stronger inductive biases and reduced shortcut reliance.
CM2 replaces verifiable outcome rewards with checklist rewards for multi-turn tool-use RL, improving 8B models by 8, 12 points on agent benchmarks using simulated environments.