Claude vs Codex · judged by Gemini
What this proves
The coding-agent debate is usually abstract. Here are three real tasks, both agents run side by side, with a third agent judging the code excerpts and bugs. No cherry-picking, no vibes.
Three real coding tasks. Claude Code and Codex each run from the same task brief in a fresh sandbox. Gemini 3 Pro scores correctness, quality, speed, and fit. See the scored excerpts, the bugs, and the verdict.