# An AI model tracker — what each new model scores, what it costs, and which to try

Recipe No. 31, Work and teams. From The know.sh Cookbook: https://know.sh/cookbook/ai-benchmarks

Five labs shipped six models this month, and each claims the crown. Your tracker puts every model’s SWE-bench Verified, GPQA Diamond and Humanity’s Last Exam scores in one table, plots score against price, and keeps the team’s call beside the evidence. A weekly run by your assistant adds whatever shipped since Monday.

- For: a team lead keeping up with new AI models for the team
- You get: a monthly tracker of scores and prices, and the call on which to try
- Time: an hour to set up, then a weekly run
- Made with: Shelf, Document, Elements, Your AI assistant, Scheduled runs

## How it works

1. **Start the month’s tracker.** Ask your assistant for a document per month: a table element of models against benchmarks, a scatter of score against price, and a section per week.
2. **Schedule the Monday run.** Set a weekly scheduled task for new model releases and their published scores. When it arrives, reply *file this in know.sh*, and your assistant adds each model with its sources.
3. **Check the numbers yourself.** Open each source and check every score before it counts. Mark a lab’s own figures as unverified until an independent leaderboard or your own test agrees.
4. **Record the call.** Ask for a decisions element: which model to try, why, what else was considered, and which call it replaces. Mark the pick in the table.

## Try this prompt

Your assistant, connected to know.sh (Claude, ChatGPT or a local model):

> Using know.sh, start a document for this month on my AI models shelf. In the overview, a table element with a row per model and columns for SWE-bench Verified, GPQA Diamond, Humanity’s Last Exam, ARC-AGI-2 and blended price per million tokens, the model we trial marked as the pick. Under it, a scatter of SWE-bench Verified against price, labelled by model. Cite every score.

Your assistant, connected to know.sh (Claude, ChatGPT or a local model):

> Every Monday, find AI models released in the past week and the benchmark scores published for them, saying which come from the lab and which from an independent leaderboard. Then file them in know.sh: add each to this month’s table, scatter and timeline, with sources.

## Elements in the specimens

Written as your assistant writes them through the know.sh MCP server. Each reads as plain words until you turn elements on (Account → Elements).

Table, specimen 1:

```element table
{
  "caption": "Seven models on four benchmarks, September 2026",
  "columns": [
    {
      "label": "Model",
      "align": "left"
    },
    {
      "label": "SWE-bench",
      "align": "right",
      "numeric": true
    },
    {
      "label": "GPQA",
      "align": "right",
      "numeric": true
    },
    {
      "label": "HLE",
      "align": "right",
      "numeric": true
    },
    {
      "label": "ARC-AGI-2",
      "align": "right",
      "numeric": true
    },
    {
      "label": "$/M tok",
      "align": "right",
      "numeric": true
    }
  ],
  "rows": [
    [
      "Ridge 2, in use",
      "74.9",
      "84.2",
      "26.3",
      "24.0",
      "6.00"
    ],
    [
      "Ridge 3",
      "82.4",
      "88.1",
      "34.7",
      "41.0",
      "6.00"
    ],
    [
      "Kestrel 2.5",
      "80.6",
      "86.9",
      "31.2",
      "36.5",
      "2.40"
    ],
    [
      "Lumen 4",
      "85.1",
      "90.3",
      "41.5",
      "52.0",
      "14.00"
    ],
    [
      "Lumen 4 Mini",
      "74.6",
      "81.0",
      "22.9",
      "19.5",
      "0.90"
    ],
    [
      "Weft 72B",
      "71.3",
      "78.4",
      "18.6",
      "12.0",
      "0.45"
    ],
    [
      "Sable 1",
      "80.9",
      "87.5",
      "35.0",
      "38.0",
      "3.10"
    ]
  ],
  "pick": 2,
  "notes": "Scores are per cent. SWE-bench Verified: share of 500 human-checked GitHub issues resolved. GPQA Diamond: 198 graduate-level science questions. HLE: *Humanity’s Last Exam*, 2,500 expert-written questions, answered without tools. ARC-AGI-2: grid puzzles easy for people and hard for models. Price blends three input tokens to one output. Sable 1’s figures are the lab’s own."
}
```

Scatter, specimen 1:

```element scatter
{
  "caption": "SWE-bench Verified against price, the same seven models",
  "xLabel": "Price, $ per million tokens",
  "xMin": 0,
  "xMax": 15,
  "yLabel": "SWE-bench Verified",
  "yUnits": "%",
  "yMin": 70,
  "yMax": 86,
  "points": [
    {
      "label": "Ridge 2",
      "x": 6,
      "y": 74.9,
      "note": "What the team uses today."
    },
    {
      "label": "Ridge 3",
      "x": 6,
      "y": 82.4
    },
    {
      "label": "Kestrel 2.5",
      "x": 2.4,
      "y": 80.6,
      "note": "The pick: within 1.8 points of Ridge 3 at 40 per cent of the price."
    },
    {
      "label": "Lumen 4",
      "x": 14,
      "y": 85.1
    },
    {
      "label": "Lumen 4 Mini",
      "x": 0.9,
      "y": 74.6
    },
    {
      "label": "Weft 72B",
      "x": 0.45,
      "y": 71.3,
      "note": "Open weights; price is a hosted endpoint."
    },
    {
      "label": "Sable 1",
      "x": 3.1,
      "y": 80.9,
      "note": "The lab’s own figure, not yet independently run."
    }
  ],
  "trend": false,
  "equality": false,
  "quadrants": true,
  "xDivide": 5,
  "yDivide": 78,
  "topLeft": "Worth a trial",
  "topRight": "Best, at a price",
  "bottomLeft": "Cheap, for bulk work",
  "bottomRight": "Skip",
  "annotations": [],
  "source": "Scores from each lab’s announcement or the SWE-bench leaderboard, checked 28 September. The models and labs are invented."
}
```

Timeline, specimen 2:

```element timeline
{
  "caption": "Model releases, September 2026",
  "events": [
    {
      "date": "2026-09-02",
      "title": "Halden Ridge 3",
      "body": "The successor to what we use: 7.5 points up on SWE-bench Verified, same price."
    },
    {
      "date": "2026-09-09",
      "title": "Corvid Kestrel 2.5",
      "body": "Within 1.8 points of Ridge 3 at $2.40 per million tokens."
    },
    {
      "date": "2026-09-11",
      "title": "Our 60-issue replay",
      "body": "Kestrel 41 of 60, Ridge 3 43, Ridge 2 36, on issues from our own repository."
    },
    {
      "date": "2026-09-16",
      "title": "Meridian Lumen 4 and Lumen 4 Mini",
      "body": "Top of every table we keep, at $14.00. The Mini is priced for bulk work."
    },
    {
      "date": "2026-09-22",
      "title": "Tessel Weft 72B",
      "body": "Open weights, so it could run on our own hardware."
    },
    {
      "date": "2026-09-24",
      "title": "Brightwater Sable 1",
      "body": "A new lab; its scores are its own so far."
    }
  ]
}
```

Decisions, specimen 2:

```element decisions
{
  "caption": "Which model the team uses, as of 28 September",
  "decisions": [
    {
      "title": "Model for the code-review bot",
      "status": "decided",
      "date": "2026-09-25",
      "choice": "Trial Kestrel 2.5 on one pull request in ten for two weeks",
      "reason": "Within two issues of Ridge 3 on our own replay, at 40 per cent of the price.",
      "alternatives": [
        {
          "option": "Ridge 3"
        },
        {
          "option": "Lumen 4"
        }
      ]
    },
    {
      "title": "Upgrade to Ridge 3",
      "status": "superseded",
      "date": "2026-09-04",
      "choice": "Move from Ridge 2 to Ridge 3 at the same price",
      "reason": "The obvious step, until Kestrel matched it for less.",
      "supersededBy": "Model for the code-review bot"
    },
    {
      "title": "Hard tickets",
      "status": "decided",
      "date": "2026-09-25",
      "choice": "Lumen 4, only when the trial model fails",
      "reason": "Best on every benchmark we track, and nearly six times Kestrel’s price."
    },
    {
      "title": "Sable 1",
      "status": "open",
      "choice": "Wait for an independent SWE-bench Verified run",
      "reason": "On the lab’s own figures it ties Kestrel. Revisit when it reaches a public leaderboard."
    },
    {
      "title": "MMLU-Pro in the tracker",
      "status": "decided",
      "date": "2026-09-07",
      "choice": "Dropped",
      "reason": "Every model we track scores between 84 and 88; it no longer tells them apart."
    }
  ]
}
```
