Home › Compare
Claude Code vs Codex: the comparison lab
The lab we're building: the same task, rule file and session-carry test on every agent, each result linked to a probe you can run. Until the probes are published, the matchups below show the format, not results. No affiliate links.
Concept page — the matchups, meters, figures and answers below are illustrative until the probes behind them are published. Check each vendor for current plans and limits.
The matrix
Cells ship only where probes actually run — anything else is one more opinion listicle, and the internet is stocked.
| Matchup | Status |
|---|---|
| Codex vs Claude Code | Probe-backed |
| Claude Code vs Cursor | Partial |
| OpenCode vs Claude Code | Probe-backed |
| Claude Code vs GitHub Copilot | Registered |
| Gemini CLI vs Claude Code | Registered |
| Codex vs Cursor | Watch list |
What we measure that nobody else does
Same "never push to main" rule on each runtime's own hooks — who blocks, who asks, who times out open. Enforced / Best-effort straight from the coverage report.
Kill it mid-task, resume elsewhere. What survives? With our layer: files, memory, and rules do.
"Yes, even the expensive ones stop."