3. Getting facts right

Benchmark

A public exam every AI company sits, so the results can go in a league table. It's an eval — a fixed set of questions with known answers — but a shared one, set by outsiders, on questions that have nothing to do with your business. These are the scores quoted in the announcements.

League tables, with the usual caveat. Top of the table tells you a team is good. It doesn't tell you they'll win on your pitch, in your weather.

Why it matters: a benchmark score tells you a model is generally capable and nothing at all about your task on your data. Never pick a model on benchmarks alone — run your own golden set, which is the same idea using your questions and your correct answers. Twenty of your own real cases will tell you more than any league table.

Related: Eval · Golden set

Bence K. Csernak

Bence K. Csernak

Founder

Newsletter