AI
Models, tools and agents, tested on our own machines
Gemma 4 with thinking off: 15 runs, no empty replies and answers in half the time
With thinking off in Ollama, every run answered and the typical reply took 27.2 seconds instead of 53.8. The writing speed did not change.
- Research
Ollama’s default context length for Gemma 4 31B on a 48 GB Mac is 32,768, and memory climbs as you use it
Ollama chose 32,768 tokens on its own and listed 19 GB. The process itself grew from 22.0 GB to 33.2 GB over 15 runs at 65,536.
- Guides
How to import a local GGUF file into Ollama with a Modelfile, using Gemma 4 31B
Two commands run a GGUF file you already have. What Ollama showed after the import, and the second copy of the file it makes on disk.
- Guides
How to turn off thinking in Ollama for Gemma 4, and what it changed in our test
One setting stopped the empty replies and cut the typical wait roughly in half. Two ways to set it, how to check your model supports it, and the numbers from 15 runs.
- Reviews
Ollama vs LM Studio on an M4 Pro: the same Gemma 4 file runs at the same speed
Same Mac, same model file, same 600 token limit: 11.74 tokens a second in LM Studio and 11.50 in Ollama. The difference is in what each lets you change.
- Reviews
Gemma 4 31B QAT on an M4 Pro with 48 GB: 11.7 tokens a second, after a long think
It fits a 48 GB Mac with room to spare and writes at a steady pace. It also spends most of its tokens thinking before the reader sees a word.
- Research
Why Gemma 4 gives an empty response: we counted where its tokens go
In 18 timed runs the model spent most of its tokens on thinking nobody sees. With 600 tokens allowed, it ran out before answering in 6 runs of 15.
- Guides
How to measure tokens per second in LM Studio without the cache fooling you
Run the same prompt twice and a model looks hundreds of times quicker to start than it is. Five steps to a figure you can trust.