Skip to benchmark content
MotionBenchHosted by Baz Studio
July diagnostic pilot33 routes · diagnostic datasetJuly diagnostic pilot
openai · C

GPT-5.1

gpt-5.1-2025-11-13
Diagnostic tierB
Human assessment
Good fragment field and clean final hold.

Historical diagnostic evidence; not a durable capability claim.

GPT-5.1 — MotionBench