Skip to benchmark content
MotionBenchHosted by Baz Studio
July diagnostic pilot33 routes · diagnostic datasetJuly diagnostic pilot
openai · O+C

GPT 5.4

gpt-5.4
Diagnostic tierC
Human assessment
Valid but small and visually conservative in sampled frames.

Historical diagnostic evidence; not a durable capability claim.