Latest news
Announcements, benchmark releases, and the work behind both.
Research & Engineering•2 min read
FrontierSWE v2
34 tasks, 13 frontier models, 20-hour budgets. Our hardest coding benchmark yet.
Read more →
Leaderboard
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5.156.3% | 56.3% |
| 2 | GPT-5.632.2% | 32.2% |
| 3 | GLM-5.330.2% | 30.2% |
| 4 | Kimi K325.9% | 25.9% |
| 5 | Grok 4.625.3% | 25.3% |
| 6 | Gemini 3.7 Flash20.3% | 20.3% |
| 7 | Qwen3.8-Max15.8% | 15.8% |
| 8 | DeepSeek V4 Flash Exp14.8% | 14.8% |
| 9 | Muse Spark 1.212.0% | 12.0% |
| 10 | Inkling4.1% | 4.1% |
Research & Engineering•9 min read
FrontierSWE
Our ultra-long-horizon coding benchmark. Frontier models clear only a fraction of its tasks.
Read more →Research & Engineering•7 min read
Our Problems
An overview of the problems we’re working on at Proximal.
Read more →Company•4 min read
Announcing Proximal
We believe data is becoming one of the central research problems in AI, and no one is working on it the right way. Proximal is a research lab for data.
Read more →