Skip to benchmark contentMotionBenchHosted by Baz Studio Human assessment
July diagnostic pilot33 routes · diagnostic datasetJuly diagnostic pilot
openai · C
GPT-5.1
gpt-5.1-2025-11-13Diagnostic tierB
Good fragment field and clean final hold.
Historical diagnostic evidence; not a durable capability claim.