100,000+ videos. One benchmark for motion graphics.
Every AI model gets the same frozen prompt and brand pack. Each candidate then acts as its own director and engineer — planning, building, reviewing, and repairing the video itself — before the same fixed renderer exports it, so you can compare the finished videos, human preference, completion, time, and cost.
The production system behind MotionBench has powered 100,000+ videos for 13,000 creators.
How the benchmark works
- Same inputEvery model starts with the same prompt and frozen brand pack.
- The candidate runs the whole buildThe same model decides the story, plans the scenes, writes the code, and reviews its own work.
- Same rendererOne fixed renderer and export path turns every project into its MP4.
- Compare the MP4sWatch the full result, vote blind, then inspect time and cost.
The input stays fixed.
Changing the starting prompt would change the test. Only the candidate model building the video changes.
“Make a 10-second launch video for baz.studio. Decide the strongest product story, create your own creative brief and execution plan, then build, review, and deliver the finished video. No audio.”
July diagnostic pilot
This development snapshot proves the candidate-controlled publication path end to end. It is clearly marked non-official and cannot be promoted from synthetic evidence.
Time and raw model cost
Lower-left is faster and cheaper. Every point is named; color shows the sampled visual tier, and an open point failed a technical gate.
Scroll horizontally to inspect every labeled route →
Run the same comparison with your own prompt.
Choose two to eight models. Your results stay separate from the official benchmark.
Open Community Lab