Search
Guides

How to measure tokens per second in LM Studio without the cache fooling you

Run the same prompt twice and a model looks hundreds of times quicker to start than it is. Five steps to a figure you can trust.

Written byadmin
Read time3 min
Published3 October 2026
Filed underGuides

LM Studio reports a speed for every reply, but one easy mistake makes a model look far quicker than it is: running the same prompt twice. In our test, a model that took 46 seconds to start answering a long prompt started in a third of a second when sent the same prompt again. This guide shows how to get a figure you can trust.

What does tokens per second mean in LM Studio?

It is how fast the model writes once it has begun. LM Studio reports it beside a second figure that matters as much: the time to the first token, which is how long the model spends reading your prompt before it writes anything. A model can write quickly and still keep you waiting, so record both.

Why is the second run of the same prompt so much faster?

LM Studio keeps the prompt it has just read. If the next prompt starts with the same text, it skips the reading. We sent one 3,806 token prompt to Gemma 4 31B three times without changing it, then once more with a different first line.

Run Wait for the first token Whole reply
First time 46.1 seconds 61.9 seconds
Same prompt again 0.32 seconds 16.0 seconds
Same prompt a third time 0.11 seconds 15.9 seconds
Different first line 47.1 seconds 66.4 seconds

The repeat was 144 times quicker to start, and the third run more than 400 times quicker. The writing speed did not change: about 11.2 tokens a second in every run. Anyone who warms up with a prompt and then times the same prompt is measuring the saved copy.

How do I measure tokens per second properly in LM Studio?

  1. Load the model and note the context length it was loaded with. Speed and memory both depend on it.
  2. Write three to five prompts of the kind you really use: a question, some code, a summary.
  3. Put a different first line at the top of every run, such as a run number. A changed first line forces the model to read the whole prompt again.
  4. Run each prompt three times and write down the tokens per second and the time to the first token that LM Studio shows.
  5. Take the middle result of each, not the best one.

How do I get tokens per second from the LM Studio server?

Start the local server and send your prompt to its own address, /api/v0/chat/completions, on port 1234. The reply carries a stats block with tokens_per_second, time_to_first_token and generation_time, so a short script can run the whole test and save every figure. The addresses are listed in LM Studio’s developer documentation.

What should I write down with the result?

A speed means little without what it ran on. Record the model and its exact build, the machine and its memory, the runtime and its version, the context length, how many tokens you allowed, and the date. That is the same list we keep for every entry in the index.

What is a good tokens per second for a local model?

It depends on what you will wait for. For a model that thinks before it answers, the wait matters more than the rate. In our test of Gemma 4 31B the model wrote at 11.7 tokens a second, and a 217 word answer still took 91 seconds.