Model pricing for coding agents
Benchmarks rank models. This ranks what they cost to run in a loop — grouped by the job you’d actually hire them for.
Hardest reasoning, highest cost. Worth it for architecture, gnarly debugging, and one-shot work you cannot babysit.
| Why you'd pick it | |||||||
|---|---|---|---|---|---|---|---|
| Claude Opus 5anthropic/claude-opus-5 | The default when the task is ambiguous or the blast radius is large. Best at holding a big refactor in its head without losing the thread. | 1M | $5.00 | $25.00 | $0.500 | Full | |
| Claude Fable 5anthropic/claude-fable-5 | Priced well above its siblings — reach for it when output quality matters more than throughput. | 1M | $10.00 | $50.00 | $1.00 | — | |
| GPT 5.6 Solopenai/gpt-5.6-sol | Strong general reasoning with a very large context window. A solid alternative when you want a second opinion from outside the Anthropic family. | 1.1M | $2.00 | $10.00 | $0.200 | Partial | |
| GPT 5.6 Terraopenai/gpt-5.6-terra | Same generation as Sol with a different balance; worth A/B-ing on your own workload rather than trusting a leaderboard. | 1.1M | $2.00 | $12.00 | $0.200 | Partial | |
| GPT 5.5 Proopenai/gpt-5.5-pro | Priced for deliberation, not throughput. Use it on the one hard problem, not the loop around it. | 1M | $30.00 | $180.00 | — | — | |
| Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | Cheap for a frontier model and very long context — good for tasks that need to read a lot before writing a little. | 1M | $2.00 | $12.00 | $0.200 | Partial | |
| Grok 4.6spacexai/grok-4.6 | Competitive on price for its tier. Worth testing if your work is heavy on tool-calling. | 500K | $2.00 | $6.00 | $0.500 | — | |
| Kimi K3moonshotai/kimi-k3 | Frontier-class at a mid-tier price, though less proven in agent harnesses than the majors. | 1M | $3.00 | $15.00 | $0.300 | Full |