Skip to benchmark contentMotionBenchHosted by Baz Studio Human assessment
July diagnostic pilot33 routes · diagnostic datasetJuly diagnostic pilot
openai · O+C
GPT 5.4
gpt-5.4Diagnostic tierC
Valid but small and visually conservative in sampled frames.
Historical diagnostic evidence; not a durable capability claim.