Skip to benchmark contentMotionBenchHosted by Baz Studio Human assessment
July diagnostic pilot33 routes · diagnostic datasetJuly diagnostic pilot
openai · C
GPT-4.1 Mini FT
ft-gpt-4.1-mini-codegen-v1Diagnostic tierB
Clean centered result with a credible authored transition.
Historical diagnostic evidence; not a durable capability claim.