Aikido Security announced Altar-1, an open-weight security model made by removing 88 of the 256 routed experts in each layer of Z.AI’s GLM-5.3, on September 21, 2026. Hugging Face’s own data for the repository counts 500,825,296,352 parameters in 327.99 GB of files. We have not run it. Our test machine, an Apple M4 Pro with 48 GB of memory, has less than a sixth of the memory those files need, so every figure below about how the model behaves is Aikido’s claim and not our measurement.
Takeaway points
- Hugging Face counts 500,825,296,352 parameters in the Altar-1 repository, which keeps 168 of the 256 routed experts in each layer of GLM-5.3.
- On Aikido’s own benchmark of 32 known vulnerabilities, Altar averaged 60.4% recall against 65.6% for full-precision GLM-5.3, the announcement says.
- We have not run Altar-1: its 327.99 GB of weight files is more than the 48 GB of memory on our test machine.
What is Altar-1, and what did Aikido remove from GLM-5.3?
Altar-1 is GLM-5.3 with fewer experts and fewer bits per weight. GLM-5.3 is a mixture-of-experts model, a design in which each token is routed through a few small sub-networks instead of the whole model. Z.AI’s config file for GLM-5.3 lists 256 routed experts per layer, of which 8 are chosen per token. The config file in the Altar-1 repository lists 168 routed experts and the same 8 per token. Aikido’s announcement states the arithmetic: “Altar retains 168 of the original 256 routed experts in each backbone expert layer, removing 88, or 34.4% of them.”
The pruning method is REAP, short for router-weighted expert activation pruning, which the card credits to Cerebras Research. The paper, first posted to arXiv on October 15, 2025, describes a criterion that “considers both router gate-values and expert activation norms”. Aikido did not use the score alone. The card says each expert was ranked by its largest share of the routed work in any single domain, so that the specialists for code, rare languages and structured output are kept.
The repository also carries the list of survivors. The file reap_prune_plan.json, 59,099 bytes long, has five keys. Its keep value is 168 and its criterion is the string “massmax168”. Its source is a local path that names an AWQ INT4 checkpoint of GLM-5.3. Its layers key holds 76 entries, numbered 3 to 78, and each entry is a sorted list of 168 expert numbers between 0 and 255. The fifth key, mtp_ranking, is a note that reads “weight magnitude from” followed by another local path.
The 76 lists are all different. By our count of the file, every one of the 256 expert numbers is kept in at least one layer and none is kept in all 76. The file holds no scores, so it shows which experts were kept and not why. GLM-5.3’s config lists 78 hidden layers, of which the first 3 are dense. The earliest version of the card, dated September 7, said the prune covered “all 75 MoE layers + MTP”, which accounts for 76 entries.
How big is Altar-1 compared with GLM-5.3?
Altar-1 is 327.99 GB of files, which is 43% of the size of the GLM-5.3 download, by our arithmetic. The table sets out the numbers from Hugging Face’s data for each repository.
| Item | Z.AI GLM-5.3 | Altar-1 |
|---|---|---|
| Parameters counted by Hugging Face | 753,329,940,480 | 500,825,296,352 |
| Routed experts per layer, config file | 256 | 168 |
| Experts chosen per token, config file | 8 | 8 |
| Files in the repository | 755.66 GB | 327.99 GB |
Of the Altar-1 total, 473,820,037,120 parameters are stored as 32-bit integers, 27,005,246,464 as 16-bit floating point and 12,768 as 32-bit floating point. The config file sets a packed format with 4-bit weights for the quantised layers. The model card says only the routed experts are 4-bit, and that attention, the shared expert, the dense layers and the head stay at 16 bits. In Z.AI’s repository, Hugging Face counts 751,226,191,872 of the parameters in 8-bit floating point.
The repository lists 48 files, 39 of them weight files in the safetensors format. It is not gated. It holds no GGUF or MLX build.
The size reduction depends on the baseline. Aikido’s announcement compares Altar’s 328.0 GB with the other two forms of GLM-5.3 it describes.
| Baseline named by Aikido | Stored weights | Reduction to Altar |
|---|---|---|
| GLM-5.3 unpruned, 16-bit | 1,506.7 GB | 78.2% |
| GLM-5.3 unpruned, AWQ INT4 | 488.2 GB | 32.8% |
| Altar, pruned, 4-bit weights and 16-bit activations | 328.0 GB | not applicable |
Z.AI’s own repository is neither of those baselines. Its files come to 755.66 GB, and Hugging Face counts its weights mostly as 8-bit floating point, so the 78.2% figure is measured against a 16-bit form that is not the download. The announcement does not make the comparison with the download.
How does Altar-1 score against GLM-5.3 on Aikido’s tests?
On Aikido’s own benchmark of 32 known vulnerabilities, Altar averaged 60.4% recall per run against 65.6% for full-precision GLM-5.3, the company says. The benchmark covers “32 known vulnerabilities across 30 repositories, with three runs per case.”
| Model, as reported by Aikido | Average recall per run | Vulnerabilities found at least once |
|---|---|---|
| Altar | 60.4% | 23 of 32 |
| GLM-5.3, AWQ INT4 | 61.5% | 23 of 32 |
| GLM-5.3, full precision | 65.6% | 25 of 32 |
Aikido puts the drop from the full model at 5.2 percentage points. These are the company’s numbers from its own test. We have not repeated the benchmark and we cannot, because the 32 cases are not published.
Aikido also states the limits of the test. “It does not measure blind discovery across an entire codebase, execute exploits to validate findings, or evaluate the fix-proposal stage,” the announcement says, and the pipeline “uses other models for surrounding stages.” Neither the card nor the announcement gives a score on any public benchmark for coding, reasoning or languages.
The card gives a second measure. It reports a KL divergence of 0.506 nats against full 16-bit GLM-5.3, measured on what it calls a “sealed 25-prompt panel, full 154k vocabulary”. KL divergence scores how differently two models predict the next token, and a score of 0 means identical, the card says. The repository holds no evaluation files. For the full numbers the card points to a dataset hosted outside Aikido’s organisation, which we did not open and do not use as a source.
Aikido also says it deployed Altar to its Aikido Machine fleet after the evaluation, where “it identified a valid critical-severity vulnerability during a client production pentest.” The announcement gives no further detail of that finding.
Can Altar-1 run on a Mac with 48 GB of memory?
No, not as published, and we have not tried. The files are 327.99 GB, which is 6.83 times the memory of our test machine. The card says Altar-1 was built to be served on four Nvidia H200 cards under the vLLM engine and adds: “Requires Hopper (H100/H200).” It says 328 GB of weights on four H200 cards “leaves room for a 128k-context KV cache at production batch sizes”. Its serve command splits the model across 4 cards and sets a maximum length of 131,072 tokens. The config file keeps GLM-5.3’s 1,048,576 maximum positions.
Pruning does not remove the problem for a desktop. Only 8 experts are chosen per token, but Aikido’s announcement says of such models that “the full expert pool still has to be stored and served.” Our reading is that a model this size needs its weights held in memory somewhere, and a 48 GB machine cannot hold 328 GB. Memory also grows with context. For a sense of how much a much smaller model asks of our machine, see our test of Gemma 4 31B QAT on an M4 Pro with 48 GB, and for how memory climbs with context, Ollama’s default context length for Gemma 4 31B on a 48 GB Mac.
If a build small enough for our machine ever appears, we would measure it the way we describe in how to measure tokens per second in LM Studio without the cache fooling you. None exists in this repository.
What licence does Altar-1 carry, and where does the card disagree with the repository?
The licence field on the Altar-1 card is “other”, and the card says the model inherits the GLM-5.3 licence. The Altar-1 file listing has no licence file of its own, so a person who downloads it receives no copy of the terms. The card links to Z.AI’s repository instead.
Z.AI’s file is titled “GLM-5.3 License”. It grants the right to use, modify, distribute and sell the weights, on two conditions. The copyright and permission notice must be included “in all copies or substantial portions of the Software.” The second applies to a company that runs a model-as-a-service business and whose revenue, with its affiliates, exceeds 10 billion US dollars over any consecutive 12 months. Such a company “must pass Z.AI’s security review before using the Software or its derivative works for any commercial purpose.”
The repository and the card disagree in four other places.
| What the card or announcement says | What the repository or Hugging Face shows |
|---|---|
| Card title: “a 504B parameter Prune of GLM-5.3” | Hugging Face counts 500.8 billion parameters |
Serve command: vllm serve aikido/altar-1 |
The repository is AikidoSec/altar-1. Our request to Hugging Face’s API for aikido/altar-1 on October 5 returned a 401 error, and the request for AikidoSec/altar-1 returned the repository’s data |
| 78.2% smaller | 78.2% is measured against 16-bit GLM-5.3 at 1,506.7 GB, and Z.AI’s own download is 755.66 GB |
| Base model is GLM-5.3 | Hugging Face’s metadata lists Altar-1 as a quantised version of the cyankiwi checkpoint, not of Z.AI’s repository |
The name has also changed. Hugging Face records the repository as created on September 7, and all 39 weight files date from a commit made that day. The card’s title then began “GLM-5.3-504B”. A commit on September 15 renamed the model Aura-1, and a second commit 21 minutes later changed the name to Altar-1. The last commit, on September 20, added a section which says “some of Aikido’s products are already powered by Altar”. The announcement followed on September 21, two weeks after the weights were uploaded, and uses the name Altar without the number.
The card credits four parties: Z.AI for the base model, an account named cyankiwi for the 4-bit checkpoint, Cerebras Research for REAP, and an account named 0xSero, which the card says “performed the REAP prune”. The announcement thanks the same four and spells the second “cyanwiki”. We report this only as what the card and the announcement say. The card says the work was done on eight Nvidia RTX PRO 6000 Blackwell cards.
Have we run Altar-1, and what has Aikido not published?
We have not run Altar-1. We read the model card, the repository listing, the config files, the prune plan and Aikido’s announcement, and we did the arithmetic on the figures in them. This report carries no speed or memory measurements of our own, because the files are more than six times the memory of our test machine. Every performance figure here is Aikido’s claim.
Aikido has not published the 32 cases of its vulnerability benchmark or the 25 prompts behind the KL figure, so neither result can be reproduced from the repository. The calibration data is described and not released. The card says Altar-1 “was calibrated on cybersecurity traces, coding, tool calling, reasoning, and English”, and the announcement says “No customer data was ever involved.”
The card does not explain the gap between the 504B in its title and the 500.8 billion parameters Hugging Face counts, and it has not corrected the repository name in its serve command. Aikido has not said whether it will add a copy of Z.AI’s licence to the repository.
For scale, DASLab’s compressed build of Qwen3.8 Flash Next, another open-weight release we have reported on, comes to 58.41 GB.
Sources
Cite this page
Hill. "Aikido announces Altar-1, GLM-5.3 cut to 168 experts and 327.99 GB". Onticpost, published 5 October 2026. https://onticpost.com/aikido-altar-1-glm-5-3-pruned/
Updated 5 October 2026. This address stays the same when the piece is updated.