Good Papers

MOVEBENCH: A Benchmark for Global-Scale Wildlife Movement Forecasting

MoveBench introduces a 2.6M-location wildlife movement forecasting benchmark across 110 species and finds existing methods generalize poorly to unseen individuals and deep learning does not consistently beat simpler baselines.

Justin Kay, Shir Bar, Ellen O Aikens, Martin Becker, Francesca Cagnacci, Juliet Cohen, Scott W Forrest, Jessica Kendall-Bar, Madeleine Lucas, Macon Overcast, Meredith S Palmer, Will Rogers, Nicholas J Russo, Christian Rutz, Larissa T Beumer, Michael B Brown, Ying-Chi Chan, Sarah C Davidson, Diego E Soto, Anne G Hertel, Roland Kays, Benjamin Koger, Guram Mikaberidze, Thomas Mueller, Ruth Oliver, Thorsten Papenbrock, Robert Patchett, Jared A Stabach, Dane Taylor, Scott W Yanco, Sara Beery

Published 2026Atlanta Poster Session 4 · Thu, Dec 10, 4:30 PM–7:30 PM local time · Hall C1arXiv ↗OpenReview ↗

91%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
?1 reader voted. Vote to see how they split.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel16/20reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
AI panel?Vote to see what the 20 AI reviewers said

Abstract

Understanding and predicting wildlife movement is critical for ecology and conservation. While trajectory forecasting has advanced for human and vehicle movement, wildlife trajectories present distinct challenges: they are unconstrained in space, highly stochastic, and influenced by environmental conditions. We introduce MoveBench, the first large-scale benchmark for probabilistic wildlife movement forecasting, containing 2.6M GPS locations from 800+ individuals across 110 species in 127 countries, paired with 1.6B environmental raster tiles capturing 160 covariates known or hypothesized to influence movement. We propose a probabilistic evaluation protocol for movement trajectory forecasts, addressing limitations of point-prediction metrics for inherently stochastic phenomena. Through comprehensive empirical evaluation of four method families across multiple temporal and spatial scales, we reveal that: (1) existing predictive methods generalize better to future timepoints than to unseen individuals, (2) deep learning approaches do not consistently outperform simpler baselines, and (3) environmental covariate selection significantly impacts performance. MoveBench enables standardized evaluation of movement forecasting methods and provides a foundation for methodological advances on this ecologically important task.