Best LLMs for Coding — 2026 Rankings

Coding LLM Leaderboard

Which AI model writes the best code? We rank every major LLM — open and closed source — across software engineering, code generation, competitive programming, and agentic coding benchmarks.

Roshan Desai

Roshan Desai · Last updated: 2026-07-20

S

Claude Fable 5

GPT-5.6 Sol

Kimi K3

2.8T

A

Claude Opus 4.8

GPT-5.5

DeepSeek-V4-Pro

1.6T

Kimi K2.6

1T

GLM-5.2

753B

GPT-5.6 Terra

Claude Sonnet 5

MiniMax M3

428B

DeepSeek-V4-Flash

284B

Step-3.7-Flash

198B

Qwen3.6-27B

27B

Gemini 3.1 Pro

N/A

GPT-5.6 Luna

B

Qwen3.6-35B-A3B

35B

DeepSeek V3.2

685B

DeepSeek R1

671B

MiMo-V2-Flash

309B

Step-3.5-Flash

196B

Qwen 3.5

397B

C

MiMo-V2.5-Pro

1.02T

GPT-oss 120B

117B

Nemotron Ultra 253B

253B

D

Llama 4 Maverick

400B

Grok 3

N/A

Cost vs. Coding Performance

Which models give you the best coding performance for the price? Top-left is the sweet spot — high performance, low cost.

Top-left = best value (high performance, low cost). Models without pricing data are excluded.

Best Coding LLMs by Benchmark

How does each model perform on real-world software engineering, code generation, competitive programming, and terminal-based coding tasks? Hover any bar for details.

Best at Software Engineering

Real-world software engineering tasks (SWE-bench Verified)

Best for Code Generation

Python code generation from docstrings (HumanEval)

Best in Competitive Coding

Real-world coding problems (LiveCodeBench)

Best at Terminal Coding

Agentic terminal coding tasks (Terminal-Bench 2.1)

Hardest Software Engineering

Harder, contamination-resistant engineering tasks (SWE-bench Pro)

Coding Benchmark Scores & Pricing

Complete coding benchmark results and pricing for every model. Click any column header to sort.

Filter:

Claude Fable 5

Anthropic

1M

$10.00

$50.00

N/A

N/A

95.0

N/A

N/A

N/A

80.0

84.3

Claude Opus 4.8

Anthropic

1M

$5.00

$25.00

N/A

93.6

88.6

N/A

N/A

N/A

69.2

74.6

Claude Sonnet 5

Anthropic

1M

$2.00

$10.00

N/A

N/A

85.2

N/A

N/A

N/A

63.2

80.4

DeepSeek R1

DeepSeek

671B

128K

$0.28

$0.42

84.0

71.5

49.2

90.2

65.9

N/A

N/A

N/A

DeepSeek V3.2

DeepSeek

685B

130K

$0.28

$0.42

85.0

79.9

67.8

N/A

74.1

39.6

N/A

N/A

DeepSeek-V4-Flash

DeepSeek

284B

1M

$0.14

$0.28

86.2

88.1

79.0

N/A

91.6

56.9

N/A

N/A

DeepSeek-V4-Pro

DeepSeek

1.6T

1M

$0.43

$0.87

87.5

90.1

80.6

N/A

93.5

67.9

N/A

N/A

Gemini 3.1 Pro

Google

N/A

1M

$2.00

$12.00

85.0

91.9

78.0

93.0

81.3

56.2

N/A

N/A

GLM-5.2

Zhipu AI

753B

1M

$1.40

$4.40

N/A

91.2

N/A

N/A

N/A

N/A

N/A

81.0

GPT-5.5

OpenAI

1M

$5.00

$30.00

88.1

93.6

88.7

94.2

78.0

82.7

59.4

85.6

GPT-5.6 Luna

OpenAI

1M

$1.00

$6.00

N/A

92.3

N/A

N/A

N/A

N/A

62.7

84.7

GPT-5.6 Sol

OpenAI

1M

$5.00

$30.00

N/A

94.6

N/A

N/A

N/A

N/A

64.6

88.8

GPT-5.6 Terra

OpenAI

1M

$2.50

$15.00

N/A

92.9

N/A

N/A

N/A

N/A

63.4

87.4

GPT-oss 120B

OpenAI

117B

128K

N/A

N/A

90.0

80.9

62.4

88.3

60.0

18.7

N/A

N/A

Grok 3

xAI

N/A

131K

$3.00

$15.00

N/A

84.6

49.0

94.5

79.4

52.0

N/A

N/A

Kimi K2.6

Moonshot

1T

262K

$0.95

$4.00

N/A

90.5

80.2

N/A

89.6

66.7

N/A

N/A

Kimi K3

Moonshot

2.8T

1M

$3.00

$15.00

N/A

93.5

N/A

N/A

N/A

N/A

N/A

88.3

Llama 4 Maverick

Meta

400B

1M

N/A

N/A

80.5

69.8

N/A

62.0

43.4

N/A

N/A

N/A

MiMo-V2-Flash

Xiaomi

309B

262K

N/A

N/A

84.9

83.7

73.4

84.8

80.6

38.5

N/A

N/A

MiMo-V2.5-Pro

Xiaomi

1.02T

1M

$0.43

$0.87

68.5

66.7

N/A

N/A

39.6

N/A

N/A

N/A

MiniMax M3

MiniMax

428B

1M

$0.30

$1.20

N/A

N/A

80.5

N/A

N/A

N/A

59.0

N/A

Nemotron Ultra 253B

Nvidia

253B

128K

N/A

N/A

N/A

76.0

N/A

N/A

66.3

N/A

N/A

N/A

Qwen 3.5

Qwen

397B

262K

N/A

N/A

87.8

88.4

76.4

N/A

83.6

52.5

N/A

N/A

Qwen3.6-27B

Qwen

27B

262K

$0.60

$3.60

86.2

87.8

77.2

N/A

83.9

59.3

N/A

N/A

Qwen3.6-35B-A3B

Qwen

35B

262K

N/A

N/A

85.2

86.0

73.4

N/A

80.4

51.5

N/A

N/A

Step-3.5-Flash

Stepfun

196B

262K

$0.10

$0.30

85.8

N/A

74.4

81.1

86.4

51.0

N/A

N/A

Step-3.7-Flash

Stepfun

198B

262K

$0.20

$1.15

N/A

N/A

76.5

N/A

N/A

N/A

N/A

N/A

Compare Coding LLMs Head-to-Head

Select two models to see how they compare across all coding and reasoning benchmarks.

Model A

Model B

Claude Opus 4.8

GPT-5.5

GPQA Diamond

93.6

vs

93.6

SWE-bench Verified

88.6

vs

88.7

SWE-bench Pro

69.2

vs

59.4

Terminal-Bench 2.1

74.6

vs

85.6

Benchmarks won

1

vs

2

Try These Models in Onyx

Onyx is the open-source AI platform that lets you connect any of these LLMs to your team's docs, apps, and people.

Best LLM for Coding 2026 | AI Coding Model Rankings & Benchmarks | Onyx AI