IQuest released IQuest-Q1, a mixture-of-experts model for agentic coding, as open weights on September 28, 2026. Hugging Face counts 320,318,615,552 parameters in the repository, stored in about 650 GB of files, and the model card says an estimated 15 billion are active for each token. We have not run it. Our test machine, an Apple M4 Pro with 48 GB of memory, holds less than one thirteenth of those files, so every figure below is IQuest’s claim or a count from the repository, not a measurement of ours.
Takeaway points
- Hugging Face counts 320,318,615,552 parameters in the IQuest-Q1 repository, and the model card says an estimated 15 billion are active for each token.
- IQuest’s own chart of eight tests places IQuest-Q1 first on none of them.
- We have not run IQuest-Q1: its files come to about 650 GB, against the 48 GB of memory on our test machine.
What is IQuest-Q1, and how big is it?
IQuest-Q1 is a model with 320 billion parameters in total, of which the model card says an estimated 15 billion are activated per token. That is about 4.7% of the total. In a mixture-of-experts design each token is routed through a few small sub-networks, called experts, instead of the whole model. The card lists 88 transformer layers and 256 experts, of which 8 are chosen per token. It says IQuest built the model “for agentic coding, reasoning, and multi-step tool use.”
The repository’s config file matches the card. It names the architecture IQuestQ1ForCausalLM, with a hidden size of 3,072, 48 attention heads, 8 key-value heads, a vocabulary of 160,000 tokens and 524,288 maximum positions. Its layer lists show one full attention layer, then a block of one full attention layer and three sliding-window layers repeated 21 times, then three full attention layers. That is 88 layers, 63 of them sliding-window. The sliding window covers 4,096 tokens.
The card gives the context length as 524,288 tokens. Its setup notes for Claude Code use the model name IQuest-Q1[1m], and the card says that client-side setting “does not change 512K context limit.” A long context costs memory as it fills. Ollama’s default context length for Gemma 4 31B on a 48 GB Mac shows what that looked like for a model we did run. We did not test IQuest-Q1 at any context length.
The weights are stored at 16 bits. The Hugging Face API lists 175 safetensors shards that total 640.64 GB, the largest of them 5.34 GB. At two bytes each, 320,318,615,552 parameters come to 640.64 GB, which matches the shards to within 0.14 MB. By that arithmetic the parameter count does not cover a further 9.34 GB file in the repository’s mtp folder. That file holds the multi-token prediction module, which the card’s serving commands load as a draft model for speculative decoding.
The documents give more than one figure for that module. The card’s specification table lists the module as “2 Independent (Training) / 1 Recursive x8 (Inference)” with a sliding window of 512. The config file in the mtp folder lists 7 draft slots and the same window of 512, and the main config sets the number of MTP layers to 0. IQuest does not say how the figures relate.
The model takes text only. “This checkpoint has no native image, audio, or video input capability,” the card says. The card’s metadata lists English and Chinese as its languages.
| Figure | Value | Where it comes from |
|---|---|---|
| Parameters, total | 320,318,615,552 | Hugging Face count |
| Parameters active per token | about 15 billion | model card, an estimate |
| Layers | 88 | model card and config file |
| Experts, total and chosen per token | 256 and 8 | model card |
| Context length | 524,288 tokens | model card and config file |
| Weight shards | 175 files, 640.64 GB | repository file list |
| Whole repository | 190 files, about 650 GB | repository file list |
How much hardware does IQuest-Q1 need?
The model card names no minimum hardware: no GPU model and no memory figure. It recommends two engines for production deployment, SGLang and vLLM, and every serving command it prints for either engine sets a tensor-parallel size of 8, the setting that splits a model across eight GPUs.
The vLLM route goes through IQuest’s own plugin repository, vllm-iquest-q1, whose readme names a single vLLM commit the plugin was tested with. The card also points to prebuilt container images for both engines. Its quick start assumes a running server and calls it through an OpenAI-compatible API.
The official repository holds 190 files. Its weight files are safetensors only, and it contains no GGUF or MLX build, the formats that llama.cpp, LM Studio and Apple’s MLX read. IQuest has not published a quantised build or said whether one is planned.
The official weights do not fit on our test machine: 650 GB of files is more than 13 times its 48 GB of memory. For scale, our test of Gemma 4 31B QAT on an M4 Pro with 48 GB reached 11.7 tokens a second. That is a model in the size class a 48 GB machine can hold. IQuest-Q1 is not. We cannot say how a machine of any size would run it, because we have no measurement and IQuest gives none.
What licence does IQuest-Q1 use?
IQuest-Q1 is under a modified MIT licence, and its one change is a naming requirement for commercial products. The card’s metadata gives the licence as “other”, under the name iquest-q1. The licence file is headed “Modified MIT License” and assigns the copyright to IQuest Research.
The file says the “only modification” is that anyone who uses the software or a derivative work in “commercial products or services” must “prominently display” the name IQuest-Q1 “on the user interface of such product or service.” The rest is the standard MIT text, which allows use, copying, modification and sale.
The repository is not gated, the Hugging Face API shows, so no form or approval stands between a reader and the files. The commit log records two changes to the licence on September 28, both after the commit titled “Submission of the initial version of IQuest-Q1”. We read the final file, and the terms above are those of the final file.
How does IQuest-Q1 score against Claude Opus 5 and DeepSeek?
On IQuest’s own chart, IQuest-Q1 is first on none of eight tests. IQuest publishes its scores as a chart image on the card, with no table of numbers in the text. The chart covers eight tests and shows five models on each. The table below is our transcription of the image. Every score is IQuest’s claim, and we have not reproduced any of them.
| Test | IQuest-Q1 | Rank on the chart | Leader on the chart |
|---|---|---|---|
| DeepSWE v1.1 | 64.6 | fourth | DeepSeek-V4.1-Flash, 74.2 |
| Terminal-Bench 2.1 | 83.2 | fourth | Claude Opus 5, 89.1 |
| NL2Repo | 63.0 | second | Claude Opus 5, 75.3 |
| CyberGym | 84.5 | second, level with GLM-5.3 | DeepSeek-V4.1-Flash, 88.1 |
| JobBench | 55.7 | third | Claude Opus 5, 65.7 |
| Agents’ Last Exam | 29.6 | third | Claude Opus 5, 32.2 |
| Humanity’s Last Exam | 39.2 | third | Hy4-preview, 43.4 |
| IQuest-CLIBench | 53.7 | third | GPT-5.6 Sol, 58.5 |
The last test is IQuest’s own. “IQuest-CLIBench is our in-house benchmark for CLI user experience,” the card says. On it the chart gives Claude Opus 5 58.2, just behind GPT-5.6 Sol.
The scores come from mixed sources. “For each model, we report the publicly reported score; otherwise, we evaluate the model using the corresponding benchmark setup,” the card says. It does not say which scores are which. The card says IQuest replaced images and other multimodal inputs with placeholders, because the model cannot read them, and set time limits of six hours for CyberGym and eight hours for Terminal-Bench 2.1.
The card recommends Claude Code 2.1.140 or Codex 0.142 as the agent software. It says IQuest evaluated Agents’ Last Exam with a different version, Claude Code 2.1.258. The Codex configuration printed on the card sets the approval policy to “never” and the sandbox mode to “danger-full-access”, and a comment in the same block says: “This configuration disables approval prompts and sandboxing.”
IQuest lists limits of its own. The card says IQuest-Q1 “may overlook constraints, repeat failed attempts, or leave issues unresolved” in command-line work, and that the model “remains at an early stage, with substantial limitations in its capabilities and reliability.” We found no independent results for IQuest-Q1 in the sources we reviewed.
Who is IQuest, and what did it release before IQuest-Q1?
IQuest is a research group that has released medical and coding models since 2025. The card says only that the model was “Developed by” IQuest. The Hugging Face account, IQuestLab, displays the name IQuest and is listed as a company with 18 members. The Hugging Face API does not show it as verified. The licence file names IQuest Research, which is also the title of the website at iquestlab.com. That page’s description carries the line “The Infinite Quest for Intelligence” and its keywords list AI, large models, medical intelligence, life sciences and mathematics.
The account holds 24 models. Its first repositories, created on September 16, 2025, are two medical reasoning models, Fleming-R1 at 7B and 32B. IQuest-Coder-V1, a family of code models from 7B to 40B, followed on December 30, 2025. Hugging Face counts 39,794,283,520 parameters in IQuest-Coder-V1-40B-Instruct, so IQuest-Q1 is about eight times that size. The coder card gives a context length of 128K tokens.
The coder family’s card links to a clarification that IQuest posted on its code-hosting page on January 2, 2026. It says the first SWE-bench Verified score IQuest reported, 81.4, was revised to 76.2 after the model was found to have used git commands to read commits that post-dated the issue it was solving. The clarification says “the model exhibited patterns of shortcut learning, which was unexpected.” SWE-bench Verified is not among the eight tests on the IQuest-Q1 chart.
Have we run IQuest-Q1, and what has IQuest not published?
No, we have not run IQuest-Q1. Its official files are about 650 GB, our test machine has 48 GB of memory, and the repository has no GGUF or MLX build we could load. This report carries no speed or memory measurements of our own. What it carries is counts from the repository, quotes from the model card and licence, and our transcription of the benchmark chart.
IQuest has not published a paper, technical report or announcement for the model. The card’s links lead to a code repository that repeats the card, to the vLLM plugin, to the container images and to the SGLang project’s support for the model.
Two diagrams on the card name five training stages: pre-training, mid-training, supervised fine-tuning, a stage labelled “RL & MOPD” and model merging. They show four expert models distilled into one student. The card gives no token counts, data sources, compute or energy figures for any stage, and does not spell out MOPD.
IQuest has also not stated the hardware the model needs, published a quantised build, or said which chart scores it measured itself. It has not said how the recursive x8 in its table relates to the 7 draft slots in the config.
For a model that does fit on this machine, see Ollama vs LM Studio on an M4 Pro: the same Gemma 4 file runs at the same speed. If IQuest or anyone else releases a build that fits in 48 GB, we will measure it and report the result here.
For scale, DeepSeek V4.1 Flash, another open-weight release we have reported on, comes to 510.3 GB of weight files.
Sources
Cite this page
Hill. "IQuest releases IQuest-Q1, a 320 billion parameter open weights model, on September 28". Onticpost, published 5 October 2026. https://onticpost.com/iquest-q1-open-weights/
Updated 5 October 2026. This address stays the same when the piece is updated.