Compare
| Claude Fable 5by Anthropic | GPT-5.6 Solby OpenAI | |
|---|---|---|
| Evaluation | ||
| Core fitHigher is better | 64 | 62 |
| AA Intelligence IndexHigher is better | 60 | 59 |
| Class | Frontier | Frontier |
| Rated effort | max | max |
| Coding agents | ||
| Coding Agent IndexHigher is better | 66 | 67 |
| DeepSWEHigher is better | 66 | 69 |
| Terminal-Bench v2Higher is better | 83 | 88 |
| SWE-Atlas-QnAHigher is better | 49 | 43 |
| Coding $ / taskLower is better | $11.71 | $7.08 |
| Minutes / taskLower is better | 23.4 | 10.2 |
| Cost | ||
| AA $ / taskLower is better | $2.75 | $1.04 |
| Input $ / 1MLower is better | $10.00 | $5.00 |
| Output $ / 1MLower is better | $50.00 | $30.00 |
| Cache read $ / 1MLower is better | $1.00 | $0.50 |
Core fit is Core's dated product judgment; the AA index and task cost are published by Artificial Analysis; coding results name their harness. The better value in each measured row is emphasized — rows without two measurements name no winner.