Saikei (栽景) are Japan's planted landscapes — miniature living scenes grown whole on a tray, where nothing is cut away: the art is in the cultivation.
Where the Niwaki releases prune the tree, these models prune nothing. Each is its reference model, refined in place.
Latest release
Qwen3.8-Flash-Next Saikei v1.0 4-bit
Qwen3.8-Flash-Next, un-pruned and improved: the intact 4-bit model distilled against the 2.4-trillion-parameter Qwen3.8-2.4T-A95B. Same size, same architecture, same speed — and every axis we measure improves or holds: both perplexities, the task average, and the program-state behaviours. It loads with stock mlx-vlm and stock llama.cpp — no custom code, no fork.
- 4.22WikiText-2 perplexity, from the reference's 5.06
- 0.830task average over five tasks, from 0.759
- 0 / +7program-state items lost / gained against the reference, of 1,928
- 111.6 GBon disk, unchanged; runs in 128 GB
- this release
- the reference model
Unlike the first Saikei, nothing trades off here: C4 improves too. Perplexity over 64 × 2048-token windows, paired against the reference; tasks are arc_easy, hellaswag, piqa, winogrande and boolq, 500 examples each, with boolq (0.906 from 0.732) and winogrande (0.776 from 0.682) carrying most of the gain; the program-state suite is scored by forced choice on two held-out seeds. The Q4_K GGUF reads 4.52 on WikiText-2 under llama-perplexity, which is not comparable with the MLX protocol above.
The first Saikei
Qwen3.6-35B-A3B Saikei 8-bit
Qwen3.6-35B-A3B, un-pruned and improved: the intact 8-bit model after a short knowledge-distillation pass from the 2.4-trillion-parameter Qwen3.8-2.4T-A95B. Same size, same architecture, same speed. It loads with stock mlx_lm and stock llama.cpp — no custom code, no fork.
- 6.33WikiText-2 perplexity, from the reference's 6.61
- 0.817task average over five tasks, from 0.790
- 37.7 GBon disk, unchanged; runs comfortably in 48 GB and up
- this release
- the reference model
The trade-off, stated honestly: WikiText-2 and the task suite improve; C4 drifts up toward the distillation corpus's mix. If your use leans on broad web-crawl-style text, the reference may serve you equally well. Perplexity over 64 × 2048-token windows, paired same-day against the reference; tasks are arc_easy, hellaswag, piqa, winogrande and boolq, 500 examples each. The Q8_0 GGUF reads 6.59 on WikiText-2 and 11.64 on a C4 slice under llama-perplexity, which is not comparable with the MLX protocol above.