MotionGrounder enables multi-object motion transfer via a diffusion transformer with flow-based motion signals, object-caption alignment loss, and a new object grounding score. It outperforms baselines in multi-object controllable video generation.
The framework models visible speech via directional articulatory motions composed into 3D facial animation, improving lip articulation and realism over prior methods.
SENSE uses graph-based EEG encoding and semantic conditioning to synthesize speech from brain dynamics, outperforming baselines on acoustic and semantic metrics with minimal training subjects.