Email Bench
Evaluates language models on producing production-ready emails from real briefs, in HTML and React Email.
196 tasks · 100 HTML · 96 React Email
| Model | ||||
|---|---|---|---|---|
| 1 | GPT-6 Astra | 50.0% | $0.59 | 103,640 |
| 2 | GPT-6 Sol | 48.5% | $0.12 | 108,520 |
| 3 | Claude Opus 5.5 | 36.2% | $0.42 | 247,780 |
| 4 | GPT-6 Luna | 36.2% | $0.006 | 116,668 |
| 5 | Grok 4.7 | 33.7% | $1.15 | 1,541,825 |
| 6 | Claude Fable 5.1 | 24.0% | $0.94 | 212,287 |
| 7 | Muse Spark 1.3 | 18.4% | $0.14 | 185,367 |
Every model runs each task once through the Pi coding agent, at high reasoning effort. A run that produces no email counts as a failure. Cost and tokens are averages per task.