IFM releases K2 Horizon open weights, and its smallest model will not load in LM Studio or Ollama

IFM released six K2 Horizon open-weight models on September 1, 2026. We tried IFM's own smallest GGUF file on an M4 Pro with 48 GB, and both LM Studio and Ollama refused it as an unknown architecture.

0:00

On September 1, 2026, the Institute of Foundation Models (IFM), a research lab launched by MBZUAI, released K2 Horizon, six open-weight language models under the Apache 2.0 licence. The largest, K2-Horizon-375B-A23B, holds 379,167,159,168 parameters in 758.34 GB of weight files, according to the repository data on Hugging Face. We tried the smallest model on our test machine, an Apple M4 Pro with 48 GB of memory. IFM’s own 4-bit file, K2-Horizon-1B-Q4_K_M.gguf, failed to load in both LM Studio and Ollama on October 5, with the error unknown model architecture: 'k2-horizon'.

Takeaway points

  • IFM released six K2 Horizon models, and Hugging Face counts 379,167,159,168 parameters in the largest, K2-Horizon-375B-A23B.
  • IFM’s model cards list the technical report as not yet available.
  • We tried the smallest model on our test machine: LM Studio and Ollama both stopped with unknown model architecture: ‘k2-horizon’.

How big is each K2 Horizon model?

The six models range from 1,078,285,824 to 379,167,159,168 parameters, and the names understate what the files hold. IFM created six base repositories between 22:55 and 23:25 UTC on September 1, the Hugging Face records show. Four hold dense models: K2-Horizon-0.9B, 3.7B, 7B and 32B. Two hold mixture-of-experts models, a design in which each token is routed through a few small sub-networks instead of the whole model: K2-Horizon-MoVA-36B-A4B and K2-Horizon-375B-A23B. Every card declares the Apache 2.0 licence in its metadata and none of the repositories is gated. The file lists of the six base repositories contain no separate licence file.

IFM’s organisation page says the lab was launched in May 2025 by MBZUAI and works from Abu Dhabi, Silicon Valley and Paris. Its earlier releases on the same account include K2 in April 2024, K2-Think in September 2025 and K2-V2 in December 2025.

Repository Parameters Hugging Face counts Weight files
K2-Horizon-0.9B 1,078,285,824 2.16 GB
K2-Horizon-3.7B 5,058,255,360 10.12 GB
K2-Horizon-7B 8,999,178,240 18.00 GB
K2-Horizon-32B 34,779,304,960 69.56 GB
K2-Horizon-MoVA-36B-A4B 37,444,792,020 74.89 GB
K2-Horizon-375B-A23B 379,167,159,168 758.34 GB

A technical appendix in the 3.7B repository gives the reason for that model: “3.78B core; 5.06B including embeddings”. The names count the core and leave out the embedding tables. The smallest model carries a third label. IFM’s GGUF repository for the 0.9B model names its files K2-Horizon-1B.

IFM’s documents also give two figures for the active parameters of the sparse models. The 375B card says the model “stores 375B parameters and runs 23B per token”. The family table in the 3.7B appendix lists 26.67B active parameters per token for the same model, and 5.95B for the 36B model whose name says 4B.

The config file of the 375B model sets 61 layers, a hidden size of 6,144, 48 attention heads and a vocabulary of 250,624 tokens. It lists 192 experts with 8 used per token and 1 shared expert, and sets the first three layers as dense. The maximum position count is 524,288, which the card calls a 512K context window. Five of the six models share that context length. The 0.9B config sets 131,072 positions and a vocabulary of 64,256, and its card says the model reaches that length “with YaRN RoPE scaling” from an original 8,192. The family table in the 3.7B appendix still lists the 0.9B model at 8,192, while the appendix inside the 0.9B repository gives 131,072.

The MoVA model adds routing inside attention. Its config sets 64 value experts with 4 used per token, alongside 100 feed-forward experts with 8 used per token. The modelling code describes the layer as “MoVA attention with routed value experts and optional post-attention gate.” The card expands the name to Mixture-of-Values attention and gives no further account of the mechanism.

IFM has also published quantised copies. The 375B model has an FP8 repository of 390.56 GB and an NVFP4 repository of 229.60 GB, created on September 22. The NVFP4 card says the build “Requires NVIDIA Blackwell-generation GPUs (B-series) or newer with native NVFP4 support.”

How does K2-Horizon-375B-A23B score against GLM 5.2 and other models?

Every score on the cards is IFM’s claim, run at high reasoning effort, and we have not reproduced any of them. The cards say baseline scores for other models come from Artificial Analysis where available and from IFM’s own evaluation code otherwise.

The 375B card claims agentic results that match or beat open-weight mixture-of-experts models up to 2.6 times its size. Its table compares the model with four open models and three closed ones.

Benchmark (IFM’s figures) K2-Horizon-375B-A23B Comparison in IFM’s table
Toolathlon Verified 65.3 59.9 GLM 5.2, 71.6 Claude Sonnet5
GDPVal-AA rating 1,441 1,498 GLM 5.2
Terminal-Bench 2.1 70.2 77.9 GLM 5.2
BrowseComp 72.8 83.5 MiniMax-M3
Humanity’s Last Exam, no tools 32.0 39.0 MiniMax-M3, 41.1 GLM 5.2

GLM 5.2 is listed at 753 billion parameters. In this table the model is ahead of GLM 5.2 on Toolathlon Verified and behind it on the other four tests.

The dense 32B model trails a smaller competitor in IFM’s own table. The card gives it 36.6 on Terminal-Bench 2.1 against 79.8 for Qwen3.8-27B, and 22.5 on tau3-Banking against 48.0. That table labels the IFM column “K2-Horizon-32B-Stage1”, a name the card does not explain.

Two tables on the 3.7B and 7B cards disagree with each other. The main table gives K2-Horizon-3.7B 68.6 on SWE-bench Verified. A second table, on a training stage IFM added to stop repeated output, gives 67.6 before that stage and 65.6 after it, and the card says “after” is the released checkpoint. For the 7B model the main table says 70.6 and the second table 69.2 before and 72.4 after. The cards do not say which figure a user of the released file should expect.

Which K2 Horizon models fit in 48 GB of memory?

On file size alone, the 0.9B, 3.7B and 7B models fit in 48 GB at full precision, and the 32B and MoVA models fit only as quantised files. We have not measured memory use for any of them, so what follows compares file sizes with our test machine.

The 375B model is out of reach. Its 758.34 GB of 16-bit weights is almost 16 times the machine’s memory, and the 229.60 GB NVFP4 copy is more than four times. IFM’s card says its serving recipe was validated on one node of eight H200 GPUs, and the hardware table in the 3.7B appendix advises “at least 16x 80 GiB GPUs on lower-memory hardware.”

The 32B and MoVA repositories hold 69.56 GB and 74.89 GB, above the machine’s memory. IFM publishes GGUF repositories for five models, all but the 375B, each with six builds. The Q4_K_M files are 21.08 GB for the 32B model and 22.37 GB for the MoVA model, and the Q8_0 files are 36.97 GB and 39.83 GB. For a sense of what a dense model of about that size does on this machine, see our test of Gemma 4 31B QAT on an M4 Pro with 48 GB. It is a different model and says nothing about K2 Horizon.

Context is the other cost. IFM’s hardware table estimates the 16-bit key-value cache of the 7B model at 18.0 GiB for one request of 128K tokens, more than the 16.8 GiB it gives for the weights. These are IFM’s estimates, not ours.

Does K2 Horizon run in LM Studio or Ollama?

Not yet, in our test. On October 5, 2026, at 09:58 BST, we loaded IFM’s K2-Horizon-1B-Q4_K_M.gguf, a file of 666,184,064 bytes from the 0.9B GGUF repository, on our test machine, which was running macOS 27.0.1. The file is about 72 times smaller than the machine’s memory, so memory was not the limit.

In LM Studio, with the llama.cpp runtime version 2.47.0, the command lms load k2-horizon-0.9b --context-length 8192 --gpu max -y ended with “Failed to load model” and the cause error loading model: unknown model architecture: 'k2-horizon'. In Ollama 0.35.1, we created a model from the file with a Modelfile and asked it to reply with one word. Ollama returned a 500 error, “llama-server process has terminated”, with the same message: unknown model architecture: 'k2-horizon'. For how that Modelfile route works, see our guide to importing a GGUF file into Ollama. Ollama’s error names a llama-server process, and LM Studio’s names a llama.cpp runtime, so both failed in the same engine family. We have measured them running one Gemma 4 file at the same speed.

This is what the official card told us to expect. The GGUF cards say the files need a version of llama.cpp “containing K2 Horizon architecture support”, that a pull request to llama.cpp “is in progress”, and that IFM keeps a fork of llama.cpp, named MBZUAI-IFM, with the code until then. We checked that wording in the saved 0.9B GGUF card before relying on it. The cards add that the GGUF builds “have currently been evaluated only on non-agent tasks.”

The state of the upstream requests on October 5, as our log records it, was as follows.

Request Title as logged State on October 5
llama.cpp pull request 29535 model : add K2 Horizon dense and MoVA support Open, created September 27
llama.cpp issue 28361 Eval bug: K2-Horizon models fail to load Open, created September 4
mlx-lm issue 1876 Support for k2_horizon Request. Open, created September 10

The same log shows that the latest mlx-lm on PyPI was 0.32.0 that day, and that LM Studio offered llama.cpp runtime 2.51.0 while we had 2.47.0 installed. We did not apply the update, so we do not know whether it changes the result. The official account lists no MLX build.

How was K2 Horizon trained, and what has IFM not published?

IFM describes the training path in detail and has not published a technical report. The 375B card lists pretraining in two phases of 7.08 trillion and 8.05 trillion tokens at a sequence length of 8K. Four midtraining stages follow, which extend the sequence length to 32K, 128K and then 512K. IFM says it then trained five expert models with reinforcement learning, for knowledge work, instruction following, search, tool use and reasoning, merged them, and ran three phases of supervised fine-tuning. The 3.7B, 7B, 32B and MoVA cards each list 22.9 trillion pretraining tokens over 1,100,000 steps. The 0.9B card lists 5 trillion and says the model was then trained by distillation from teacher models for math and code, STEM, and instruction following.

IFM has published the path as well as the result. The cards index branches for intermediate and final checkpoints of each stage, and the commit history of the 375B repository shows those branches preserved on September 23. The training code is in a GitHub repository named xllm, created on September 2 under the Apache 2.0 licence. The cards point to a dataset named TxT360-v2, which Hugging Face lists under a CC BY 4.0 licence.

The cards list the technical report as “Not yet available” with an expected date of “End of September 2026”. That entry was unchanged on October 5 in cards last updated on September 28 and October 1. The card metadata names two datasets, K2-Horizon-Pretrain-Data and K2-Horizon-Midtrain-Data, and both addresses returned an authorisation error to us on October 5. The 375B card says token counts for the reinforcement learning stage “are not reported here”. None of the documents gives the hardware, time or energy used in training. The appendix says the family “is not optimized for moderation, refusal behavior, or production assistant safety” and is “documented primarily with English-language evaluations”.

Have we run K2 Horizon?

We have tried to, and the smallest model did not load. That is the whole of our own data on K2 Horizon: two failed loads of one file on one machine on one morning, with no speed or memory figure.

What we did not do matters as much. We did not run IFM’s own llama.cpp fork, which the cards name as the way to run the GGUF files, so we cannot say whether the model works there or how fast it is. We did not test the 7B or 36B GGUF files, although they are downloaded, and the 0.9B result does not prove those files fail the same way. Our reading is that they would, since the error names the architecture and not the size, but we have not shown it. We did not test MLX. We did not apply the newer LM Studio runtime. We did not run the model in any way that produced text, so every benchmark in this report is IFM’s claim and none is ours. We tested nothing about the 32B model, the 375B model or the quantised copies.

For scale, Xiaomi’s MiMo V2.6 family, another open-weight release we have reported on, has a largest model of 573.46 GB of weight files.

Sources

  1. IFM, K2-Horizon-375B-A23B model card
  2. IFM, K2-Horizon-0.9B-GGUF model card
Cite this page

Hill. "IFM releases K2 Horizon open weights, and its smallest model will not load in LM Studio or Ollama". Onticpost, published 5 October 2026. https://onticpost.com/ifm-k2-horizon-open-weights/

Updated 5 October 2026. This address stays the same when the piece is updated.

Share this