LM Studio
Latest
Ollama vs LM Studio on an M4 Pro: the same Gemma 4 file runs at the same speed
Same Mac, same model file, same 600 token limit: 11.74 tokens a second in LM Studio and 11.50 in Ollama. The difference is in what each lets you change.
- Reviews
Gemma 4 31B QAT on an M4 Pro with 48 GB: 11.7 tokens a second, after a long think
It fits a 48 GB Mac with room to spare and writes at a steady pace. It also spends most of its tokens thinking before the reader sees a word.
- Research
Why Gemma 4 gives an empty response: we counted where its tokens go
In 18 timed runs the model spent most of its tokens on thinking nobody sees. With 600 tokens allowed, it ran out before answering in 6 runs of 15.
- Guides
How to measure tokens per second in LM Studio without the cache fooling you
Run the same prompt twice and a model looks hundreds of times quicker to start than it is. Five steps to a figure you can trust.