Best LLMs for Coding — 2026 Rankings
Coding LLM Leaderboard
Which AI model writes the best code? We rank every major LLM — open and closed source — across software engineering, code generation, competitive programming, and agentic coding benchmarks.
Roshan Desai · Last updated: 2026-07-20
S
Claude Fable 5
GPT-5.6 Sol
Kimi K3
2.8T
A
Claude Opus 4.8
GPT-5.5
DeepSeek-V4-Pro
1.6T
Kimi K2.6
1T
GLM-5.2
753B
GPT-5.6 Terra
Claude Sonnet 5
MiniMax M3
428B
DeepSeek-V4-Flash
284B
Step-3.7-Flash
198B
Qwen3.6-27B
27B
Gemini 3.1 Pro
N/A
GPT-5.6 Luna
B
Qwen3.6-35B-A3B
35B
DeepSeek V3.2
685B
DeepSeek R1
671B
MiMo-V2-Flash
309B
Step-3.5-Flash
196B
Qwen 3.5
397B
C
MiMo-V2.5-Pro
1.02T
GPT-oss 120B
117B
Nemotron Ultra 253B
253B
D
Llama 4 Maverick
400B
Grok 3
N/A
Cost vs. Coding Performance
Which models give you the best coding performance for the price? Top-left is the sweet spot — high performance, low cost.
Top-left = best value (high performance, low cost). Models without pricing data are excluded.
Best Coding LLMs by Benchmark
How does each model perform on real-world software engineering, code generation, competitive programming, and terminal-based coding tasks? Hover any bar for details.
Best at Software Engineering
Real-world software engineering tasks (SWE-bench Verified)
Best for Code Generation
Python code generation from docstrings (HumanEval)
Best in Competitive Coding
Real-world coding problems (LiveCodeBench)
Best at Terminal Coding
Agentic terminal coding tasks (Terminal-Bench 2.1)
Hardest Software Engineering
Harder, contamination-resistant engineering tasks (SWE-bench Pro)
Coding Benchmark Scores & Pricing
Complete coding benchmark results and pricing for every model. Click any column header to sort.
Filter:
Claude Fable 5 Anthropic | 1M | $10.00 | $50.00 | N/A | N/A | 95.0 | N/A | N/A | N/A | 80.0 | 84.3 | |
Claude Opus 4.8 Anthropic | 1M | $5.00 | $25.00 | N/A | 93.6 | 88.6 | N/A | N/A | N/A | 69.2 | 74.6 | |
Claude Sonnet 5 Anthropic | 1M | $2.00 | $10.00 | N/A | N/A | 85.2 | N/A | N/A | N/A | 63.2 | 80.4 | |
DeepSeek R1 DeepSeek | 671B | 128K | $0.28 | $0.42 | 84.0 | 71.5 | 49.2 | 90.2 | 65.9 | N/A | N/A | N/A |
DeepSeek V3.2 DeepSeek | 685B | 130K | $0.28 | $0.42 | 85.0 | 79.9 | 67.8 | N/A | 74.1 | 39.6 | N/A | N/A |
DeepSeek-V4-Flash DeepSeek | 284B | 1M | $0.14 | $0.28 | 86.2 | 88.1 | 79.0 | N/A | 91.6 | 56.9 | N/A | N/A |
DeepSeek-V4-Pro DeepSeek | 1.6T | 1M | $0.43 | $0.87 | 87.5 | 90.1 | 80.6 | N/A | 93.5 | 67.9 | N/A | N/A |
Gemini 3.1 Pro | N/A | 1M | $2.00 | $12.00 | 85.0 | 91.9 | 78.0 | 93.0 | 81.3 | 56.2 | N/A | N/A |
GLM-5.2 Zhipu AI | 753B | 1M | $1.40 | $4.40 | N/A | 91.2 | N/A | N/A | N/A | N/A | N/A | 81.0 |
GPT-5.5 OpenAI | 1M | $5.00 | $30.00 | 88.1 | 93.6 | 88.7 | 94.2 | 78.0 | 82.7 | 59.4 | 85.6 | |
GPT-5.6 Luna OpenAI | 1M | $1.00 | $6.00 | N/A | 92.3 | N/A | N/A | N/A | N/A | 62.7 | 84.7 | |
GPT-5.6 Sol OpenAI | 1M | $5.00 | $30.00 | N/A | 94.6 | N/A | N/A | N/A | N/A | 64.6 | 88.8 | |
GPT-5.6 Terra OpenAI | 1M | $2.50 | $15.00 | N/A | 92.9 | N/A | N/A | N/A | N/A | 63.4 | 87.4 | |
GPT-oss 120B OpenAI | 117B | 128K | N/A | N/A | 90.0 | 80.9 | 62.4 | 88.3 | 60.0 | 18.7 | N/A | N/A |
Grok 3 xAI | N/A | 131K | $3.00 | $15.00 | N/A | 84.6 | 49.0 | 94.5 | 79.4 | 52.0 | N/A | N/A |
Kimi K2.6 Moonshot | 1T | 262K | $0.95 | $4.00 | N/A | 90.5 | 80.2 | N/A | 89.6 | 66.7 | N/A | N/A |
Kimi K3 Moonshot | 2.8T | 1M | $3.00 | $15.00 | N/A | 93.5 | N/A | N/A | N/A | N/A | N/A | 88.3 |
Llama 4 Maverick Meta | 400B | 1M | N/A | N/A | 80.5 | 69.8 | N/A | 62.0 | 43.4 | N/A | N/A | N/A |
MiMo-V2-Flash Xiaomi | 309B | 262K | N/A | N/A | 84.9 | 83.7 | 73.4 | 84.8 | 80.6 | 38.5 | N/A | N/A |
MiMo-V2.5-Pro Xiaomi | 1.02T | 1M | $0.43 | $0.87 | 68.5 | 66.7 | N/A | N/A | 39.6 | N/A | N/A | N/A |
MiniMax M3 MiniMax | 428B | 1M | $0.30 | $1.20 | N/A | N/A | 80.5 | N/A | N/A | N/A | 59.0 | N/A |
Nemotron Ultra 253B Nvidia | 253B | 128K | N/A | N/A | N/A | 76.0 | N/A | N/A | 66.3 | N/A | N/A | N/A |
Qwen 3.5 Qwen | 397B | 262K | N/A | N/A | 87.8 | 88.4 | 76.4 | N/A | 83.6 | 52.5 | N/A | N/A |
Qwen3.6-27B Qwen | 27B | 262K | $0.60 | $3.60 | 86.2 | 87.8 | 77.2 | N/A | 83.9 | 59.3 | N/A | N/A |
Qwen3.6-35B-A3B Qwen | 35B | 262K | N/A | N/A | 85.2 | 86.0 | 73.4 | N/A | 80.4 | 51.5 | N/A | N/A |
Step-3.5-Flash Stepfun | 196B | 262K | $0.10 | $0.30 | 85.8 | N/A | 74.4 | 81.1 | 86.4 | 51.0 | N/A | N/A |
Step-3.7-Flash Stepfun | 198B | 262K | $0.20 | $1.15 | N/A | N/A | 76.5 | N/A | N/A | N/A | N/A | N/A |
Compare Coding LLMs Head-to-Head
Select two models to see how they compare across all coding and reasoning benchmarks.
Model A
Model B
Claude Opus 4.8
GPT-5.5
GPQA Diamond
93.6
vs
93.6
SWE-bench Verified
88.6
vs
88.7
SWE-bench Pro
69.2
vs
59.4
Terminal-Bench 2.1
74.6
vs
85.6
Benchmarks won
1
vs
2
Try These Models in Onyx
Onyx is the open-source AI platform that lets you connect any of these LLMs to your team's docs, apps, and people.