Every figure we publish comes from a run we can repeat. This page says how, so you can check it or run your own.
How do we time a model?
We write several prompts of the kind people really use: an explanation, some code, a summary, a list and a worked sum. We run each three times and report the middle result, not the best one. Every prompt starts with a different first line, so a program cannot reuse a prompt it has already read. Why that matters is explained in our guide to timing a model.
What do we record with each result?
| What | Why |
|---|---|
| The model and its exact build | Two files with the same name can differ |
| The machine and its memory | Speed depends on both |
| The program and its version | The same file can run differently in each |
| The context length and token limit | Both change speed, memory and whether you get an answer |
| The date | Programs and models change |
How do we label a figure?
Every figure in the index carries one of three labels. Measured by us means we ran it. Reproduced means we repeated someone else’s result. Community means readers sent it, and we show how many machines are behind it.
What do we do when a maker gives a figure?
We show it beside ours, in a table headed Maker says and We found. If the maker gives no figure, we say so.
How is a score decided?
A score is out of five and is stated at the top of each review with the verdict. The review says what it was based on.