Niwaki (庭木) are Japan's garden trees, sculpted by meticulous pruning so that every branch serves the form of the whole.

This project applies that spirit to Mixture-of-Experts language models: open weights pruned to fit the memory you have, measured against the unpruned model, and released for MLX and llama.cpp.

Engraving of a niwaki: a Japanese garden pine with a twisted trunk and cloud-pruned foliage.

Latest release

Qwen3.6 30B · 27B · 24B Niwaki v3.0 · per-expert precision

Qwen3.6-35B-A3B with each of its 10,240 routed experts stored at its own precision — 4, 3 or 2 bits — or removed, in three sizes that all run in 24 GB of unified memory. The 30B holds the 8-bit reference's task average at 3.4× less expert storage: perplexity 3–8% above the reference's, program-state tracking at its level, and ten major languages within 4–21% of it.

  • 10.1 GBof expert storage in the 30B, from 34.2 GB in the 8-bit reference
  • 99.7%of the reference's task average (0.794 against 0.796)
  • −12%WikiText-2 perplexity against v2.3, at 11% fewer expert bytes (6.98 against 7.95)
  • 10 languageskept within 4–21% of the reference's perplexity, English to Hindi
WikiText-2 perplexitylower is better
  1. Reference, 8-bit · 34.2 GB of experts6.76
  2. v3.0 · 30B · 10.1 GB6.98
  3. v3.0 · 27B · 8.9 GB7.20
  4. v3.0 · 24B · 7.3 GB7.80
  5. Unpruned, experts at 2-bit, retrained · 10.1 GB7.66
  6. v2.3 · 4-bit · 11.3 GB7.95
  7. v2.3 · 2-bit · 7.3 GB9.02
Task average, five taskshigher is better
  1. Reference, 8-bit · 34.2 GB of experts0.796
  2. v3.0 · 30B · 10.1 GB0.794
  3. v3.0 · 27B · 8.9 GB0.791
  4. v3.0 · 24B · 7.3 GB0.785
  5. Unpruned, experts at 2-bit, retrained · 10.1 GB0.793
  6. v2.3 · 4-bit · 11.3 GB0.772
  7. v2.3 · 2-bit · 7.3 GB0.760
Ten languages, perplexity against the referencebest to worst · lower is better
  1. v3.0 · 30B1.04–1.21
  2. v3.0 · 27B1.06–1.24
  3. v3.0 · 24B1.11–1.34
  4. v2.3 · 4-bit1.14–1.23
  5. v2.3 · 2-bit1.33–1.55
Two-step arithmetic, generated answershigher is better
  1. Reference, 8-bit0.63–0.68
  2. v3.0 · 30B0.53–0.62
  3. v3.0 · 27B0.58–0.62
  4. v3.0 · 24B0.48–0.57
  5. Unpruned, experts at 2-bit, retrained0.53–0.60
  6. v2.3 · 4-bit0.97–0.98
  7. v2.3 · 2-bit0.92–0.93
  • Niwaki v3.0
  • earlier Niwaki generations
  • the unpruned model

Perplexity over WikiText-2's full test set in 2048-token windows; task average over the complete arc_easy, hellaswag, piqa, winogrande and boolq test sets; languages on FLORES-200 devtest (English, Spanish, French, German, Russian, Chinese, Japanese, Korean, Arabic and Hindi), 1.00 meaning the reference's perplexity; arithmetic ranges over two evaluation seeds. The retrained 2-bit row is the unpruned model with its routed experts at 2-bit and the same recovery training, at the 30B's exact bytes. Read it as a split on one family: v3.0 is ahead of v2.3 on everything here except two-step arithmetic in generation, where v2.3's layout stays above the reference — pick v2.3 for exact multi-step arithmetic, v3.0 for everything else. Languages beyond these ten are not protected the same way.

In llama.cpp

GGUF builds · WikiText-2 perplexity, 512-token windowslower is better
  1. v3.0 30B · 14.1 GB7.25
  2. v3.0 27B · 12.8 GB7.52
  3. v3.0 24B · 10.8 GB8.18
  4. Unpruned · UD-IQ3_S · 13.7 GB7.39
  5. Unpruned · UD-IQ3_XXS · 13.2 GB7.37
  6. Unpruned · UD-Q2_K_XL · 12.3 GB7.48
  7. v2.3 · Q4_K_M · 14.0 GB8.31
  8. v2.3 · UD-Q3K · 11.2 GB8.39

Each GGUF file reads within 0.5% of its MLX model on identical tokens, in English and in each of the ten languages. On an M4 Max with Metal (llama-bench, 512-token prompt, 128 generated tokens) the v3.0 files run at 936–978 tokens/s on the prompt and 63–64 tokens/s in decode, against 82 for v2.3's UD-Q3K build; each v3.0 layer runs its experts in two precision tiers, which costs speed against v2.3's single-precision files. The unpruned model's dynamic quants of similar size decoded at 41–43 tokens/s in an earlier session. The files need the neopolita-llama.cpp fork at tag v3.0 or later: stock llama.cpp cannot represent per-expert precision yet.

Earlier generations

  • v2.3 · Qwen3.6 23B-A3B · 4-bit 11.3 GB of experts · perplexity 7.95 · task average 0.772 MLX ↗ · GGUF ↗
  • v2.3 · Qwen3.6 23B-A3B · 2-bit 7.3 GB of experts · perplexity 9.02 · task average 0.760 MLX ↗ · GGUF ↗
  • v2.2 · Qwen3.6 23B-A3B 11.3 GB of experts · perplexity 8.14 · task average 0.770 MLX ↗ · GGUF ↗
  • v2.1 · Qwen3.6 23B-A3B 11.3 GB of experts · perplexity 8.62 · task average 0.759 MLX ↗ · GGUF ↗
  • v2 · Qwen3.6 19B-A3B 9.1 GB of experts · perplexity 9.42 · task average 0.742 MLX ↗ · GGUF ↗
  • v1 · Qwen3.6 27B-A3B 7.5 GB of experts · perplexity 9.89 · task average 0.769 MLX ↗ · GGUF ↗

From Qwen3.8-Flash-Next

Qwen3.8-Flash-Next 119B-A5B Niwaki v2.4 · 3-bit

Qwen3.8-Flash-Next pruned from about 177B to 119B parameters and stored at 3-bit: 44.8 GB on disk, for 64 GB Macs. It keeps the unpruned model's task average, and the program-state behaviours we test — variable tracking, two-step arithmetic, multi-hop lookups — stay at the unpruned model's level.

  • 44.8 GBon disk, from 111.5 GB for the 4-bit reference
  • 119B / 5.6Btotal / active parameters, from ~177B / ~6.7B
  • 100%of the unpruned model's task average (0.762 against 0.759)
  • −4 / +7program-state items lost / gained against the unpruned model, of 1,928
Size on disk, GBlower is better
  1. Unpruned, 4-bit reference111.5
  2. Unpruned, best 2-bit cast55.3
  3. Niwaki v2.4 · 119B44.8
  4. Niwaki v2.4 · 102B37.5
  5. First generation · 113B42.2
  6. First generation · 99B36.7
WikiText-2 perplexitylower is better
  1. Unpruned, 4-bit reference5.06
  2. Unpruned, best 2-bit cast6.35
  3. Niwaki v2.4 · 119B6.36
  4. Niwaki v2.4 · 102B7.84
  5. First generation · 113B7.08
  6. First generation · 99B7.88
Task average, five taskshigher is better
  1. Unpruned, 4-bit reference0.759
  2. Unpruned, best 2-bit cast0.761
  3. Niwaki v2.4 · 119B0.762
  4. Niwaki v2.4 · 102B0.738
  5. First generation · 113B0.757
  6. First generation · 99B0.733
Variable tracking past distractor lineshigher is better
  1. Unpruned, 4-bit reference0.96
  2. Niwaki v2.4 · 119B0.96
  3. Niwaki v2.4 · 102B0.96
  4. First generation · 113B (one seed)0.92
  5. First generation · 99B (one seed)0.68
  • Niwaki v2.4
  • earlier Niwaki generations
  • the unpruned model

Perplexity over 64 × 2048-token windows; task average over arc_easy, hellaswag, piqa, winogrande and boolq, zero-shot, 500 examples each, paired against the reference on identical examples; program-state suite scored by forced choice on two held-out seeds. The best 2-bit cast of the unpruned model (routed experts and n-gram table at 2-bit) reads the same perplexity and task average at 10.5 GB more. Protocols and the full tables are on the model cards.

In llama.cpp

GGUF builds · WikiText-2 perplexity, 512-token windowslower is better
  1. Niwaki v2.4 119B · UD-Q4K · 64.3 GB6.88
  2. Niwaki v2.4 119B · UD-Q3K · 51.0 GB7.04
  3. Niwaki v2.4 102B · UD-Q4K · 55.9 GB8.57
  4. Niwaki v2.4 102B · UD-Q3K · 43.9 GB8.75
  5. First generation 113B · UD-Q4K · 62.8 GB7.70
  6. First generation 113B · UD-Q3K · 49.1 GB7.90
  7. First generation 99B · UD-Q3K · 42.5 GB8.79

On an M4 Max with Metal the 119B's UD-Q3K file runs at 557 tokens/s on a 512-token prompt and 30 tokens/s in decode. The GGUF files need the neopolita-llama.cpp fork: stock llama.cpp cannot represent this model's layout yet.

For 48 GB Macs: 102B-A5B

The same release at the first-generation 99B's size: 37.5 GB on disk. Perplexity and task average match that model; what changes is function. The variable tracking the 99B lost is back at the unpruned model's level — 0.96 past distractor lines, from 0.68 — and two-step arithmetic sits at 0.87 against the unpruned model's 0.91.

  • 37.5 GBon disk, from 111.5 GB for the 4-bit reference
  • 102B / 5.3Btotal / active parameters, from ~177B / ~6.7B
  • 97%of the unpruned model's task average (0.738 against 0.759)
  • −15 / +7program-state items lost / gained against the unpruned model, of 1,928

First generation

  • Qwen3.8-Flash-Next 113B-A5B · 3-bit 42.2 GB · perplexity 7.08 · task average 0.757 MLX ↗ · GGUF ↗
  • Qwen3.8-Flash-Next 99B-A5B · 3-bit 36.7 GB, fits a 48 GB Mac · perplexity 7.88 · task average 0.733 MLX ↗ · GGUF ↗

Technical paper

A paper with the full method and measurements is coming soon.

Until then, every number on this page comes from the model cards, which carry their evaluation protocols and the full tables. About the paper