Niwaki (庭木) are Japan's garden trees, sculpted by meticulous pruning so that every branch serves the form of the whole.
This project applies that spirit to Mixture-of-Experts language models: open weights pruned to fit the memory you have, measured against the unpruned model, and released for MLX and llama.cpp.
Latest release
Qwen3.6 30B · 27B · 24B Niwaki v3.0 · per-expert precision
Qwen3.6-35B-A3B with each of its 10,240 routed experts stored at its own precision — 4, 3 or 2 bits — or removed, in three sizes that all run in 24 GB of unified memory. The 30B holds the 8-bit reference's task average at 3.4× less expert storage: perplexity 3–8% above the reference's, program-state tracking at its level, and ten major languages within 4–21% of it.
- 10.1 GBof expert storage in the 30B, from 34.2 GB in the 8-bit reference
- 99.7%of the reference's task average (0.794 against 0.796)
- −12%WikiText-2 perplexity against v2.3, at 11% fewer expert bytes (6.98 against 7.95)
- 10 languageskept within 4–21% of the reference's perplexity, English to Hindi
- Niwaki v3.0
- earlier Niwaki generations
- the unpruned model
Perplexity over WikiText-2's full test set in 2048-token windows; task average over the complete arc_easy, hellaswag, piqa, winogrande and boolq test sets; languages on FLORES-200 devtest (English, Spanish, French, German, Russian, Chinese, Japanese, Korean, Arabic and Hindi), 1.00 meaning the reference's perplexity; arithmetic ranges over two evaluation seeds. The retrained 2-bit row is the unpruned model with its routed experts at 2-bit and the same recovery training, at the 30B's exact bytes. Read it as a split on one family: v3.0 is ahead of v2.3 on everything here except two-step arithmetic in generation, where v2.3's layout stays above the reference — pick v2.3 for exact multi-step arithmetic, v3.0 for everything else. Languages beyond these ten are not protected the same way.
In llama.cpp
Each GGUF file reads within 0.5% of its MLX model on identical tokens, in English and in each of the ten languages. On an M4 Max with Metal (llama-bench, 512-token prompt, 128 generated tokens) the v3.0 files run at 936–978 tokens/s on the prompt and 63–64 tokens/s in decode, against 82 for v2.3's UD-Q3K build; each v3.0 layer runs its experts in two precision tiers, which costs speed against v2.3's single-precision files. The unpruned model's dynamic quants of similar size decoded at 41–43 tokens/s in an earlier session. The files need the neopolita-llama.cpp fork at tag v3.0 or later: stock llama.cpp cannot represent per-expert precision yet.
- 30B · MLX ↗13.6 GB · perplexity 6.98
- 27B · MLX ↗12.4 GB · perplexity 7.20
- 24B · MLX ↗10.8 GB · perplexity 7.80
Earlier generations
- v2.3 · Qwen3.6 23B-A3B · 4-bit 11.3 GB of experts · perplexity 7.95 · task average 0.772 MLX ↗ · GGUF ↗
- v2.3 · Qwen3.6 23B-A3B · 2-bit 7.3 GB of experts · perplexity 9.02 · task average 0.760 MLX ↗ · GGUF ↗
- v2.2 · Qwen3.6 23B-A3B 11.3 GB of experts · perplexity 8.14 · task average 0.770 MLX ↗ · GGUF ↗
- v2.1 · Qwen3.6 23B-A3B 11.3 GB of experts · perplexity 8.62 · task average 0.759 MLX ↗ · GGUF ↗
- v2 · Qwen3.6 19B-A3B 9.1 GB of experts · perplexity 9.42 · task average 0.742 MLX ↗ · GGUF ↗
- v1 · Qwen3.6 27B-A3B 7.5 GB of experts · perplexity 9.89 · task average 0.769 MLX ↗ · GGUF ↗
From Qwen3.8-Flash-Next
Qwen3.8-Flash-Next 119B-A5B Niwaki v2.4 · 3-bit
Qwen3.8-Flash-Next pruned from about 177B to 119B parameters and stored at 3-bit: 44.8 GB on disk, for 64 GB Macs. It keeps the unpruned model's task average, and the program-state behaviours we test — variable tracking, two-step arithmetic, multi-hop lookups — stay at the unpruned model's level.
- 44.8 GBon disk, from 111.5 GB for the 4-bit reference
- 119B / 5.6Btotal / active parameters, from ~177B / ~6.7B
- 100%of the unpruned model's task average (0.762 against 0.759)
- −4 / +7program-state items lost / gained against the unpruned model, of 1,928
- Niwaki v2.4
- earlier Niwaki generations
- the unpruned model
Perplexity over 64 × 2048-token windows; task average over arc_easy, hellaswag, piqa, winogrande and boolq, zero-shot, 500 examples each, paired against the reference on identical examples; program-state suite scored by forced choice on two held-out seeds. The best 2-bit cast of the unpruned model (routed experts and n-gram table at 2-bit) reads the same perplexity and task average at 10.5 GB more. Protocols and the full tables are on the model cards.
In llama.cpp
On an M4 Max with Metal the 119B's UD-Q3K file runs at 557 tokens/s on a 512-token prompt and 30 tokens/s in decode. The GGUF files need the neopolita-llama.cpp fork: stock llama.cpp cannot represent this model's layout yet.
For 48 GB Macs: 102B-A5B
The same release at the first-generation 99B's size: 37.5 GB on disk. Perplexity and task average match that model; what changes is function. The variable tracking the 99B lost is back at the unpruned model's level — 0.96 past distractor lines, from 0.68 — and two-step arithmetic sits at 0.87 against the unpruned model's 0.91.
- 37.5 GBon disk, from 111.5 GB for the 4-bit reference
- 102B / 5.3Btotal / active parameters, from ~177B / ~6.7B
- 97%of the unpruned model's task average (0.738 against 0.759)
- −15 / +7program-state items lost / gained against the unpruned model, of 1,928
First generation
Technical paper
A paper with the full method and measurements is coming soon.
Until then, every number on this page comes from the model cards, which carry their evaluation protocols and the full tables. About the paper