How to import a local GGUF file into Ollama with a Modelfile, using Gemma 4 31B
Two commands run a GGUF file you already have. What Ollama showed after the import, and the second copy of the file it makes on disk.
If you already have a GGUF file on your Mac, you can run it in Ollama without downloading it again. Write a one line Modelfile that points at the file, then run ollama create. We did this with the 17.65 GB Gemma 4 31B QAT Q4_0 file from Google’s model card, and it loaded first time.
How do I import a GGUF file into Ollama?
- Find the full path to the file, such as
/Users/you/models/gemma-4-31B-it-QAT-Q4_0.gguf. - In an empty folder, make a file called
Modelfilecontaining one line:FROM /Users/you/models/gemma-4-31B-it-QAT-Q4_0.gguf - Run
ollama create gemma4-31b-qat-q4 -f Modelfile. The name is yours to choose. - Run
ollama run gemma4-31b-qat-q4to chat with it, or send requests to the local server.
The server must be running first. If ollama create says it cannot connect, start it with ollama serve.
How do I check the import worked?
Run ollama show gemma4-31b-qat-q4. Ours reported:
| Field | What Ollama showed |
|---|---|
| Architecture | gemma4 |
| Parameters | 30.7B |
| Context length | 262144 |
| Embedding length | 5376 |
| Quantization | Q4_0 |
| Capabilities | completion, tools, thinking |
| Stop token | <turn|> |
The capabilities and the stop token were filled in for us from the file, so we wrote no template. The line to look at is Capabilities. It is what tells you the model can think, which matters for turning thinking off.
Does Ollama copy the model file?
Yes. After the import, Ollama’s own store at ~/.ollama/models/blobs held a 17,651,000,768 byte file, the same size as the original, and the folder measured 16 GiB. You now have two copies on disk, so check you have the space before importing a large model.
Why import a file instead of pulling a model?
To run exactly the same weights in two programs. We used one file in LM Studio and in Ollama, so any difference in speed or behaviour came from the software and not from the model.
What context length will an imported model use?
Not the 262,144 that ollama show lists. That is the most the model supports. On our 48 GB Mac, Ollama chose 32,768 on its own. The reasoning and the memory it used are in our test of Ollama’s default context.
How we tested
Ollama 0.32.6 on a Mac with an M4 Pro chip and 48 GB of memory, macOS 27.0.1. The steps above are the ones we ran. We used the local server, not ollama run, for the timed tests. The method is on how we test.
Sources
Cite this page
"How to import a local GGUF file into Ollama with a Modelfile, using Gemma 4 31B". onticpost.com, published 4 October 2026. https://onticpost.com/import-gguf-ollama-modelfile-gemma-4/
Updated 4 October 2026. This address stays the same when the piece is updated.