Trending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#4Unlimited-OCRbaidu⬇2237kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#7Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#8krea2-identity-editconradlocke⬇0kTrending Model:#9Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#10OvisOCR2ATH-MaaS⬇17kTrending Model:#1Inklingthinkingmachines⬇16kTrending Model:#2Ternary-Bonsai-27B-ggufprism-ml⬇432kTrending Model:#3Bonsai-27B-ggufprism-ml⬇1405kTrending Model:#4Unlimited-OCRbaidu⬇2237kTrending Model:#5GLM-5.2zai-org⬇545kTrending Model:#6Qwythos-9B-Claude-Mythos-5-1M-GGUFempero-ai⬇2133kTrending Model:#7Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUFDavidAU⬇63kTrending Model:#8krea2-identity-editconradlocke⬇0kTrending Model:#9Qwen3.6-35B-A3B-Uncensored-HauhauCS-AggressiveHauhauCS⬇1998kTrending Model:#10OvisOCR2ATH-MaaS⬇17k

Allenai Unleashes Molmo-motion To Forecast Future Object Movements

Floating glowing trajectory map consists of iridescent glass texture.

Molmo-motion is a new vision-language model that forecasts 3D point trajectories based on natural-language action instructions. Given a short video history, some 2D query points, and an action description, it predicts where those points will move in 3D space over a future period. This specific checkpoint uses three history frames to predict roughly two seconds of motion at 15 frames per second.

Allenai developed this tool to help systems anticipate how objects move, enabling better planning and action reasoning. To train it, the team created a large corpus of action-described, object-grounded 3D point trajectories from over one million unconstrained videos. They also built a human-verified benchmark spanning 111 object categories and 61 motion types to evaluate the model.

Features and target users

Key Features
  • Model predicts 3D point trajectories accurately.
  • Works with natural-language action instructions.
  • Supports autoregressive coordinate prediction easily.
  • Transfers well to robot manipulation tasks.

Researchers focused on trajectory prediction and motion forecasting will find this model highly useful for their studies. It also serves as a strong starting point for downstream finetuning in applications like robotics. Creators working on motion-guided video generation can leverage the predicted trajectories to produce more realistic object motion.

Developer notes and limitations

The developers note that MolmoMotion is a research model intended for educational use and should be validated before driving any downstream actuated systems. Users running the software locally must ensure their GPU drivers match the required CUDA runtime to avoid compatibility issues. Code and trained model weights are available under the Apache 2.0 license, though some datasets carry their own upstream restrictions.

"MolmoMotion is a 4B vision-language model that forecasts 3D point trajectories under natural-language action instructions." Source: Reddit