Latest news

Announcements, benchmark releases, and the work behind both.

Research & Engineering2 min read

FrontierSWE v2

34 tasks, 13 frontier models, 20-hour budgets. Our hardest coding benchmark yet.

Read more →
Leaderboard
#ModelScore
1
Claude Fable 5.156.3%
56.3%
2
GPT-5.632.2%
32.2%
3
GLM-5.330.2%
30.2%
4
Kimi K325.9%
25.9%
5
Grok 4.625.3%
25.3%
6
Gemini 3.7 Flash20.3%
20.3%
7
Qwen3.8-Max15.8%
15.8%
8
DeepSeek V4 Flash Exp14.8%
14.8%
9
Muse Spark 1.212.0%
12.0%
10
Inkling4.1%
4.1%