know.

CookbookWork and teamsNo. 31

An AI model trackerwhat each new model scores, what it costs, and which to try

Recipe No. 31 · Work and teams

For: A team lead keeping up with new AI models for the team

You get: A monthly tracker of scores and prices, and the call on which to try

Time: An hour to set up, then a weekly run

See it in know.sh

The month at a glance: seven invented models on real benchmarks, the pick marked, and score against price in four quadrants.

AI modelsNo. 9

September 2026seven models, one to try

4 sections, 1,280 words, about 6 minutes, revised Monday 28 September.

Every model the team could use this month, scored on the benchmarks we track and priced per million tokens. The pick is the one we trial next; the call and its reasons are in Week 39.

Sections

  1. 1Week 36: Halden ships Ridge 3Observation
  2. 2Week 37: Kestrel 2.5 and our 60-issue replayEvidence, keyKestrel resolved 41 of our own 60 issues; Ridge 3 resolved 43.
  3. 3Week 38: Lumen 4, and a Mini for bulk workObservation
  4. 4Week 39: Sable 1 lands, and the call on KestrelDecision, key

The week’s section: the month’s releases on a timeline, then the call in a decisions register, the superseded upgrade kept beside it.

AI modelsSeptember 2026

4of 4

Week 39: Sable 1 lands, and the call on Kestrel

Decision, key section, 6 sources, 520 words

What shipped. Six models in four weeks: one a week from the three labs we know, and two newcomers.

The call.

Claude filing Monday’s scheduled run into the tracker: the new model in the table, the scatter and the timeline, with its sources.

Claude, connected to know.sh

You

File this morning’s scheduled model run in know.sh: add Brightwater Sable 1 to the September tracker’s table and scatter, and to the Week 39 timeline. Its scores are the lab’s own, so say so, and add an open decision for it.

search Found September 2026 — seven models, one to try, AI models, No. 9

document_update_overview Added Sable 1 to the table and the scatter in September 2026: 7 models

section_update Added 24 September to the timeline and an open decision in September 2026, 4

Claude

Filed, with the announcement and model card as sources. Sable 1 edges Kestrel 2.5 on SWE-bench Verified, 80.9 to 80.6, but only on its own figures, so I left the pick on Kestrel and noted that in the table. The run found no independent result yet; next Monday’s will look again.

How it works

Five labs shipped six models this month, and each claims the crown. Your tracker puts every model’s SWE-bench Verified, GPQA Diamond and Humanity’s Last Exam scores in one table, plots score against price, and keeps the team’s call beside the evidence. A weekly run by your assistant adds whatever shipped since Monday.

  1. Start the month’s tracker

    Ask your assistant for a document per month: a table element of models against benchmarks, a scatter of score against price, and a section per week.

  2. Schedule the Monday run

    Set a weekly scheduled task for new model releases and their published scores. When it arrives, reply file this in know.sh, and your assistant adds each model with its sources.

  3. Check the numbers yourself

    Open each source and check every score before it counts. Mark a lab’s own figures as unverified until an independent leaderboard or your own test agrees.

  4. Record the call

    Ask for a decisions element: which model to try, why, what else was considered, and which call it replaces. Mark the pick in the table.

Try this prompt

Your assistant, connected to know.sh (Claude, ChatGPT or a local model):

Using know.sh, start a document for this month on my AI models shelf. In the overview, a table element with a row per model and columns for SWE-bench Verified, GPQA Diamond, Humanity’s Last Exam, ARC-AGI-2 and blended price per million tokens, the model we trial marked as the pick. Under it, a scatter of SWE-bench Verified against price, labelled by model. Cite every score.

Your assistant, connected to know.sh (Claude, ChatGPT or a local model):

Every Monday, find AI models released in the past week and the benchmark scores published for them, saying which come from the lab and which from an independent leaderboard. Then file them in know.sh: add each to this month’s table, scatter and timeline, with sources.

Made with

ShelfDocumentElementsYour AI assistantScheduled runs