Engraving of a saikei, a planted tray landscape: pines, a rock peak, a stone lantern and a small waterfall over still water, under the word Saikei.

Saikei (栽景) are Japan's planted landscapes — miniature living scenes grown whole on a tray, where nothing is cut away: the art is in the cultivation.

Where the Niwaki releases prune the tree, these models prune nothing. Each is its reference model, refined in place.

Latest release

Qwen3.8-Flash-Next Saikei v1.0 4-bit

Qwen3.8-Flash-Next, un-pruned and improved: the intact 4-bit model distilled against the 2.4-trillion-parameter Qwen3.8-2.4T-A95B. Same size, same architecture, same speed — and every axis we measure improves or holds: both perplexities, the task average, and the program-state behaviours. It loads with stock mlx-vlm and stock llama.cpp — no custom code, no fork.

  • 4.22WikiText-2 perplexity, from the reference's 5.06
  • 0.830task average over five tasks, from 0.759
  • 0 / +7program-state items lost / gained against the reference, of 1,928
  • 111.6 GBon disk, unchanged; runs in 128 GB
WikiText-2 perplexitylower is better
  1. Reference, 4-bit5.06
  2. Saikei v1.04.22
Task average, five taskshigher is better
  1. Reference, 4-bit0.759
  2. Saikei v1.00.830
C4 perplexitylower is better
  1. Reference, 4-bit10.97
  2. Saikei v1.010.29
Two-step arithmetichigher is better
  1. Reference, 4-bit0.91
  2. Saikei v1.00.94
  • this release
  • the reference model

Unlike the first Saikei, nothing trades off here: C4 improves too. Perplexity over 64 × 2048-token windows, paired against the reference; tasks are arc_easy, hellaswag, piqa, winogrande and boolq, 500 examples each, with boolq (0.906 from 0.732) and winogrande (0.776 from 0.682) carrying most of the gain; the program-state suite is scored by forced choice on two held-out seeds. The Q4_K GGUF reads 4.52 on WikiText-2 under llama-perplexity, which is not comparable with the MLX protocol above.

The first Saikei

Qwen3.6-35B-A3B Saikei 8-bit

Qwen3.6-35B-A3B, un-pruned and improved: the intact 8-bit model after a short knowledge-distillation pass from the 2.4-trillion-parameter Qwen3.8-2.4T-A95B. Same size, same architecture, same speed. It loads with stock mlx_lm and stock llama.cpp — no custom code, no fork.

  • 6.33WikiText-2 perplexity, from the reference's 6.61
  • 0.817task average over five tasks, from 0.790
  • 37.7 GBon disk, unchanged; runs comfortably in 48 GB and up
WikiText-2 perplexitylower is better
  1. Reference, 8-bit6.61
  2. Saikei6.33
Task average, five taskshigher is better
  1. Reference, 8-bit0.790
  2. Saikei0.817
C4 perplexitylower is better
  1. Reference, 8-bit10.62
  2. Saikei11.30
  • this release
  • the reference model

The trade-off, stated honestly: WikiText-2 and the task suite improve; C4 drifts up toward the distillation corpus's mix. If your use leans on broad web-crawl-style text, the reference may serve you equally well. Perplexity over 64 × 2048-token windows, paired same-day against the reference; tasks are arc_easy, hellaswag, piqa, winogrande and boolq, 500 examples each. The Q8_0 GGUF reads 6.59 on WikiText-2 and 11.64 on a C4 slice under llama-perplexity, which is not comparable with the MLX protocol above.