CookbookWork and teamsNo. 31
An AI model trackerwhat each new model scores, what it costs, and which to try
See it in know.sh
AI modelsNo. 9
September 2026seven models, one to try
4 sections, 1,280 words, about 6 minutes, revised Monday 28 September.
Every model the team could use this month, scored on the benchmarks we track and priced per million tokens. The pick is the one we trial next; the call and its reasons are in Week 39.
Seven models on four benchmarks, September 2026
| Model | SWE-bench | GPQA | HLE | ARC-AGI-2 | $/M tok |
|---|---|---|---|---|---|
| Ridge 2, in use | 74.9 | 84.2 | 26.3 | 24.0 | 6.00 |
| Ridge 3 | 82.4 | 88.1 | 34.7 | 41.0 | 6.00 |
| ✓Kestrel 2.5 | 80.6 | 86.9 | 31.2 | 36.5 | 2.40 |
| Lumen 4 | 85.1 | 90.3 | 41.5 | 52.0 | 14.00 |
| Lumen 4 Mini | 74.6 | 81.0 | 22.9 | 19.5 | 0.90 |
| Weft 72B | 71.3 | 78.4 | 18.6 | 12.0 | 0.45 |
| Sable 1 | 80.9 | 87.5 | 35.0 | 38.0 | 3.10 |
Scores are per cent. SWE-bench Verified: share of 500 human-checked GitHub issues resolved. GPQA Diamond: 198 graduate-level science questions. HLE: Humanity’s Last Exam, 2,500 expert-written questions, answered without tools. ARC-AGI-2: grid puzzles easy for people and hard for models. Price blends three input tokens to one output. Sable 1’s figures are the lab’s own.
Seven models on four benchmarks, September 2026
| Model | SWE-bench | GPQA | HLE | ARC-AGI-2 | $/M tok |
|---|---|---|---|---|---|
| Ridge 2, in use | 74.9 | 84.2 | 26.3 | 24.0 | 6.00 |
| Ridge 3 | 82.4 | 88.1 | 34.7 | 41.0 | 6.00 |
| ✓ Kestrel 2.5 | 80.6 | 86.9 | 31.2 | 36.5 | 2.40 |
| Lumen 4 | 85.1 | 90.3 | 41.5 | 52.0 | 14.00 |
| Lumen 4 Mini | 74.6 | 81.0 | 22.9 | 19.5 | 0.90 |
| Weft 72B | 71.3 | 78.4 | 18.6 | 12.0 | 0.45 |
| Sable 1 | 80.9 | 87.5 | 35.0 | 38.0 | 3.10 |
Scores are per cent. SWE-bench Verified: share of 500 human-checked GitHub issues resolved. GPQA Diamond: 198 graduate-level science questions. HLE: Humanity’s Last Exam, 2,500 expert-written questions, answered without tools. ARC-AGI-2: grid puzzles easy for people and hard for models. Price blends three input tokens to one output. Sable 1’s figures are the lab’s own.
```element table
{
"caption": "Seven models on four benchmarks, September 2026",
"columns": [
{
"label": "Model",
"align": "left"
},
{
"label": "SWE-bench",
"align": "right",
"numeric": true
},
{
"label": "GPQA",
"align": "right",
"numeric": true
},
{
"label": "HLE",
"align": "right",
"numeric": true
},
{
"label": "ARC-AGI-2",
"align": "right",
"numeric": true
},
{
"label": "$/M tok",
"align": "right",
"numeric": true
}
],
"rows": [
[
"Ridge 2, in use",
"74.9",
"84.2",
"26.3",
"24.0",
"6.00"
],
[
"Ridge 3",
"82.4",
"88.1",
"34.7",
"41.0",
"6.00"
],
[
"Kestrel 2.5",
"80.6",
"86.9",
"31.2",
"36.5",
"2.40"
],
[
"Lumen 4",
"85.1",
"90.3",
"41.5",
"52.0",
"14.00"
],
[
"Lumen 4 Mini",
"74.6",
"81.0",
"22.9",
"19.5",
"0.90"
],
[
"Weft 72B",
"71.3",
"78.4",
"18.6",
"12.0",
"0.45"
],
[
"Sable 1",
"80.9",
"87.5",
"35.0",
"38.0",
"3.10"
]
],
"pick": 2,
"notes": "Scores are per cent. SWE-bench Verified: share of 500 human-checked GitHub issues resolved. GPQA Diamond: 198 graduate-level science questions. HLE: *Humanity’s Last Exam*, 2,500 expert-written questions, answered without tools. ARC-AGI-2: grid puzzles easy for people and hard for models. Price blends three input tokens to one output. Sable 1’s figures are the lab’s own."
}
```SWE-bench Verified against price, the same seven models
| Point | Price, $ per million tokens | SWE-bench Verified | Quadrant |
|---|---|---|---|
| Ridge 2 | 6 | 74.9% | Skip |
| Ridge 3 | 6 | 82.4% | Best, at a price |
| Kestrel 2.5 | 2.4 | 80.6% | Worth a trial |
| Lumen 4 | 14 | 85.1% | Best, at a price |
| Lumen 4 Mini | 0.9 | 74.6% | Cheap, for bulk work |
| Weft 72B | 0.45 | 71.3% | Cheap, for bulk work |
| Sable 1 | 3.1 | 80.9% | Worth a trial |
Scores from each lab’s announcement or the SWE-bench leaderboard, checked 28 September. The models and labs are invented.
SWE-bench Verified against price, the same seven models
Price, $ per million tokens is drawn from 0 to 15.
SWE-bench Verified is drawn from 70% to 86%.
Divided into quadrants at Price, $ per million tokens 5 and SWE-bench Verified 78%: top left, Worth a trial; top right, Best, at a price; bottom left, Cheap, for bulk work; bottom right, Skip.
| Point | Price, $ per million tokens | SWE-bench Verified | Quadrant |
|---|---|---|---|
| Ridge 2 | 6 | 74.9% | Skip |
| Ridge 3 | 6 | 82.4% | Best, at a price |
| Kestrel 2.5 | 2.4 | 80.6% | Worth a trial |
| Lumen 4 | 14 | 85.1% | Best, at a price |
| Lumen 4 Mini | 0.9 | 74.6% | Cheap, for bulk work |
| Weft 72B | 0.45 | 71.3% | Cheap, for bulk work |
| Sable 1 | 3.1 | 80.9% | Worth a trial |
Ridge 2: What the team uses today.
Kestrel 2.5: The pick: within 1.8 points of Ridge 3 at 40 per cent of the price.
Weft 72B: Open weights; price is a hosted endpoint.
Sable 1: The lab’s own figure, not yet independently run.
Source: Scores from each lab’s announcement or the SWE-bench leaderboard, checked 28 September. The models and labs are invented.
```element scatter
{
"caption": "SWE-bench Verified against price, the same seven models",
"xLabel": "Price, $ per million tokens",
"xMin": 0,
"xMax": 15,
"yLabel": "SWE-bench Verified",
"yUnits": "%",
"yMin": 70,
"yMax": 86,
"points": [
{
"label": "Ridge 2",
"x": 6,
"y": 74.9,
"note": "What the team uses today."
},
{
"label": "Ridge 3",
"x": 6,
"y": 82.4
},
{
"label": "Kestrel 2.5",
"x": 2.4,
"y": 80.6,
"note": "The pick: within 1.8 points of Ridge 3 at 40 per cent of the price."
},
{
"label": "Lumen 4",
"x": 14,
"y": 85.1
},
{
"label": "Lumen 4 Mini",
"x": 0.9,
"y": 74.6
},
{
"label": "Weft 72B",
"x": 0.45,
"y": 71.3,
"note": "Open weights; price is a hosted endpoint."
},
{
"label": "Sable 1",
"x": 3.1,
"y": 80.9,
"note": "The lab’s own figure, not yet independently run."
}
],
"trend": false,
"equality": false,
"quadrants": true,
"xDivide": 5,
"yDivide": 78,
"topLeft": "Worth a trial",
"topRight": "Best, at a price",
"bottomLeft": "Cheap, for bulk work",
"bottomRight": "Skip",
"annotations": [],
"source": "Scores from each lab’s announcement or the SWE-bench leaderboard, checked 28 September. The models and labs are invented."
}
```Sections
- 1Week 36: Halden ships Ridge 3Observation
- 2Week 37: Kestrel 2.5 and our 60-issue replayEvidence, keyKestrel resolved 41 of our own 60 issues; Ridge 3 resolved 43.
- 3Week 38: Lumen 4, and a Mini for bulk workObservation
- 4Week 39: Sable 1 lands, and the call on KestrelDecision, key
AI modelsSeptember 2026
4of 4
Week 39: Sable 1 lands, and the call on Kestrel
Decision, key section, 6 sources, 520 words
What shipped. Six models in four weeks: one a week from the three labs we know, and two newcomers.
Model releases, September 2026
Halden Ridge 3
The successor to what we use: 7.5 points up on SWE-bench Verified, same price.
Corvid Kestrel 2.5
Within 1.8 points of Ridge 3 at $2.40 per million tokens.
Our 60-issue replay
Kestrel 41 of 60, Ridge 3 43, Ridge 2 36, on issues from our own repository.
Meridian Lumen 4 and Lumen 4 Mini
Top of every table we keep, at $14.00. The Mini is priced for bulk work.
Tessel Weft 72B
Open weights, so it could run on our own hardware.
Brightwater Sable 1
A new lab; its scores are its own so far.
Model releases, September 2026
- 2 September 2026 — Halden Ridge 3
The successor to what we use: 7.5 points up on SWE-bench Verified, same price.
- 9 September 2026 — Corvid Kestrel 2.5
Within 1.8 points of Ridge 3 at $2.40 per million tokens.
- 11 September 2026 — Our 60-issue replay
Kestrel 41 of 60, Ridge 3 43, Ridge 2 36, on issues from our own repository.
- 16 September 2026 — Meridian Lumen 4 and Lumen 4 Mini
Top of every table we keep, at $14.00. The Mini is priced for bulk work.
- 22 September 2026 — Tessel Weft 72B
Open weights, so it could run on our own hardware.
- 24 September 2026 — Brightwater Sable 1
A new lab; its scores are its own so far.
```element timeline
{
"caption": "Model releases, September 2026",
"events": [
{
"date": "2026-09-02",
"title": "Halden Ridge 3",
"body": "The successor to what we use: 7.5 points up on SWE-bench Verified, same price."
},
{
"date": "2026-09-09",
"title": "Corvid Kestrel 2.5",
"body": "Within 1.8 points of Ridge 3 at $2.40 per million tokens."
},
{
"date": "2026-09-11",
"title": "Our 60-issue replay",
"body": "Kestrel 41 of 60, Ridge 3 43, Ridge 2 36, on issues from our own repository."
},
{
"date": "2026-09-16",
"title": "Meridian Lumen 4 and Lumen 4 Mini",
"body": "Top of every table we keep, at $14.00. The Mini is priced for bulk work."
},
{
"date": "2026-09-22",
"title": "Tessel Weft 72B",
"body": "Open weights, so it could run on our own hardware."
},
{
"date": "2026-09-24",
"title": "Brightwater Sable 1",
"body": "A new lab; its scores are its own so far."
}
]
}
```The call.
Which model the team uses, as of 28 September
Model for the code-review bot
Trial Kestrel 2.5 on one pull request in ten for two weeks
Within two issues of Ridge 3 on our own replay, at 40 per cent of the price.
Also considered Ridge 3 · Lumen 4
Upgrade to Ridge 3
Move from Ridge 2 to Ridge 3 at the same price
Replaced by Model for the code-review bot
The obvious step, until Kestrel matched it for less.
Hard tickets
Lumen 4, only when the trial model fails
Best on every benchmark we track, and nearly six times Kestrel’s price.
Sable 1
Wait for an independent SWE-bench Verified run
On the lab’s own figures it ties Kestrel. Revisit when it reaches a public leaderboard.
MMLU-Pro in the tracker
Dropped
Every model we track scores between 84 and 88; it no longer tells them apart.
Which model the team uses, as of 28 September
Model for the code-review bot
✓ Decided 25 September
Choice: Trial Kestrel 2.5 on one pull request in ten for two weeks
Reason: Within two issues of Ridge 3 on our own replay, at 40 per cent of the price.
Also considered: Ridge 3 · Lumen 4
Upgrade to Ridge 3
⊘ Superseded 4 September
Choice: Move from Ridge 2 to Ridge 3 at the same price
Replaced by: Model for the code-review bot
Reason: The obvious step, until Kestrel matched it for less.
Hard tickets
✓ Decided 25 September
Choice: Lumen 4, only when the trial model fails
Reason: Best on every benchmark we track, and nearly six times Kestrel’s price.
Sable 1
○ Open
Choice: Wait for an independent SWE-bench Verified run
Reason: On the lab’s own figures it ties Kestrel. Revisit when it reaches a public leaderboard.
MMLU-Pro in the tracker
✓ Decided 7 September
Choice: Dropped
Reason: Every model we track scores between 84 and 88; it no longer tells them apart.
```element decisions
{
"caption": "Which model the team uses, as of 28 September",
"decisions": [
{
"title": "Model for the code-review bot",
"status": "decided",
"date": "2026-09-25",
"choice": "Trial Kestrel 2.5 on one pull request in ten for two weeks",
"reason": "Within two issues of Ridge 3 on our own replay, at 40 per cent of the price.",
"alternatives": [
{
"option": "Ridge 3"
},
{
"option": "Lumen 4"
}
]
},
{
"title": "Upgrade to Ridge 3",
"status": "superseded",
"date": "2026-09-04",
"choice": "Move from Ridge 2 to Ridge 3 at the same price",
"reason": "The obvious step, until Kestrel matched it for less.",
"supersededBy": "Model for the code-review bot"
},
{
"title": "Hard tickets",
"status": "decided",
"date": "2026-09-25",
"choice": "Lumen 4, only when the trial model fails",
"reason": "Best on every benchmark we track, and nearly six times Kestrel’s price."
},
{
"title": "Sable 1",
"status": "open",
"choice": "Wait for an independent SWE-bench Verified run",
"reason": "On the lab’s own figures it ties Kestrel. Revisit when it reaches a public leaderboard."
},
{
"title": "MMLU-Pro in the tracker",
"status": "decided",
"date": "2026-09-07",
"choice": "Dropped",
"reason": "Every model we track scores between 84 and 88; it no longer tells them apart."
}
]
}
```How it works
Five labs shipped six models this month, and each claims the crown. Your tracker puts every model’s SWE-bench Verified, GPQA Diamond and Humanity’s Last Exam scores in one table, plots score against price, and keeps the team’s call beside the evidence. A weekly run by your assistant adds whatever shipped since Monday.