Tokens per second
Latest
Ollama vs LM Studio on an M4 Pro: the same Gemma 4 file runs at the same speed
Same Mac, same model file, same 600 token limit: 11.74 tokens a second in LM Studio and 11.50 in Ollama. The difference is in what each lets you change.
- Reviews
Gemma 4 31B QAT on an M4 Pro with 48 GB: 11.7 tokens a second, after a long think
It fits a 48 GB Mac with room to spare and writes at a steady pace. It also spends most of its tokens thinking before the reader sees a word.
- Guides
How to measure tokens per second in LM Studio without the cache fooling you
Run the same prompt twice and a model looks hundreds of times quicker to start than it is. Five steps to a figure you can trust.