Instructions to use JohnZhan/MotionInsight-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use JohnZhan/MotionInsight-8B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("JohnZhan/MotionInsight-8B") model = AutoModelForMultimodalLM.from_pretrained("JohnZhan/MotionInsight-8B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
MotionInsight: Diagnosing Object Motion
Deficiencies in Generated Videos
EMNLP 2026 Findings
Jiahao Zhan, Yongrui Ma, Qunliang Xing, Xuanyu Zhang,
Jingqi Tong, Junlin Li, Li Zhang, Shijie Zhao, Tianfan Xue
Code · arXiv · Quick start · Citation
MotionInsight-8B
MotionInsight diagnoses object motion in generated videos using RGB frames, object tracks, and camera motion. It produces reasoning and three scores from 1 (poor) to 5 (excellent):
| Dimension | Output field |
|---|---|
| Object consistency | structural_stability |
| Motion continuity | motion_coherence |
| Physical plausibility | physical_plausibility |
The model is based on Qwen3-VL-8B-Instruct and uses four BF16 Safetensors shards (approximately 17.7 GB).
Quick start
Use Python 3.10 and the custom Qwen3-VL patch from the code repository:
git clone --recurse-submodules https://github.com/JohnZhan2023/MotionInsight.git
cd MotionInsight
python3.10 -m venv .venv-model
source .venv-model/bin/activate
python -m pip install --upgrade pip
python -m pip install huggingface-hub==0.36.2
hf download JohnZhan/MotionInsight-8B --local-dir checkpoints/MotionInsight-8B
python -m pip install torch==2.11.0 torchvision==0.26.0 \
--index-url https://download.pytorch.org/whl/cu126
python -m pip install -r checkpoints/MotionInsight-8B/requirements/inference.txt
python scripts/patch_transformers.py
After extracting motion features, run:
python inference.py \
--model_name_or_path checkpoints/MotionInsight-8B \
--video_path /path/to/video.mp4 \
--object_motion_path /path/to/object_motion.pt \
--camera_motion_path /path/to/camera_motion.pt \
--target "tennis ball" \
--output_path outputs/prediction.jsonl
Example answer format:
<thinking>Diagnostic reasoning about the target object's motion.</thinking>
<answer>{"structural_stability": 3.0, "physical_plausibility": 4.5, "motion_coherence": 3.5}</answer>
GRPO
The code repository includes a GRPO fine-tuning script:
python -m pip install -r checkpoints/MotionInsight-8B/requirements/train.txt
ATTN_IMPLEMENTATION=sdpa bash training/run_grpo.sh \
checkpoints/MotionInsight-8B /path/to/train.jsonl outputs/motioninsight-grpo
See the repository README for the input format and preprocessing setup.
License
Model weights use Apache-2.0. Preprocessing dependencies retain their own licenses; see NOTICE.
Citation
@misc{zhan2026motioninsightdiagnosingobjectmotion,
title = {MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos},
author = {Jiahao Zhan and Yongrui Ma and Qunliang Xing and Xuanyu Zhang and Jingqi Tong and Junlin Li and Li zhang and Shijie Zhao and Tianfan Xue},
year = {2026},
eprint = {2609.37030},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.37030},
}
- Downloads last month
- 74
Model tree for JohnZhan/MotionInsight-8B
Base model
Qwen/Qwen3-VL-8B-Instruct