Search
Guides

How to import a local GGUF file into Ollama with a Modelfile, using Gemma 4 31B

Two commands run a GGUF file you already have. What Ollama showed after the import, and the second copy of the file it makes on disk.

Written byadmin
Read time2 min
Published4 October 2026
Filed underAI, Guides

If you already have a GGUF file on your Mac, you can run it in Ollama without downloading it again. Write a one line Modelfile that points at the file, then run ollama create. We did this with the 17.65 GB Gemma 4 31B QAT Q4_0 file from Google’s model card, and it loaded first time.

How do I import a GGUF file into Ollama?

  1. Find the full path to the file, such as /Users/you/models/gemma-4-31B-it-QAT-Q4_0.gguf.
  2. In an empty folder, make a file called Modelfile containing one line: FROM /Users/you/models/gemma-4-31B-it-QAT-Q4_0.gguf
  3. Run ollama create gemma4-31b-qat-q4 -f Modelfile. The name is yours to choose.
  4. Run ollama run gemma4-31b-qat-q4 to chat with it, or send requests to the local server.

The server must be running first. If ollama create says it cannot connect, start it with ollama serve.

How do I check the import worked?

Run ollama show gemma4-31b-qat-q4. Ours reported:

Field What Ollama showed
Architecture gemma4
Parameters 30.7B
Context length 262144
Embedding length 5376
Quantization Q4_0
Capabilities completion, tools, thinking
Stop token <turn|>

The capabilities and the stop token were filled in for us from the file, so we wrote no template. The line to look at is Capabilities. It is what tells you the model can think, which matters for turning thinking off.

Does Ollama copy the model file?

Yes. After the import, Ollama’s own store at ~/.ollama/models/blobs held a 17,651,000,768 byte file, the same size as the original, and the folder measured 16 GiB. You now have two copies on disk, so check you have the space before importing a large model.

Why import a file instead of pulling a model?

To run exactly the same weights in two programs. We used one file in LM Studio and in Ollama, so any difference in speed or behaviour came from the software and not from the model.

What context length will an imported model use?

Not the 262,144 that ollama show lists. That is the most the model supports. On our 48 GB Mac, Ollama chose 32,768 on its own. The reasoning and the memory it used are in our test of Ollama’s default context.

How we tested

Ollama 0.32.6 on a Mac with an M4 Pro chip and 48 GB of memory, macOS 27.0.1. The steps above are the ones we ran. We used the local server, not ollama run, for the timed tests. The method is on how we test.

More in Guides