Every metric here is pulled from our live benchmark runner. No cherry-picking. No marketing claims. Raw JSON →
41 tasks across 4 coding domains. Each task independently verified by our automated evaluation pipeline. SWE-bench framework →
32 unique test queries across 8 categories, run against all 4 model tiers. Raw results →
Every competitor number sourced from public benchmarks. No estimates, no marketing.
All results are machine-generated and reproducible.
Live endpoint — no auth, no paywall.
https://ai.empire325marketing.com/v1/benchmarks/public.json
Get API Access →