Engraving of a dojo: a wooden training hall with a tiled roof and paper screens, stepping stones leading to its veranda, flanked by garden pines and a stone lantern.

Dojo

A dōjō (道場) is the place of the way: the hall where a craft is practised until it becomes one's own.

Niwaki is one long project. This page keeps the others from the same bench: model merging, documents folded into weights, a diffusion language model trained from scratch, Spanish legal models, and adapters of a few hundred kilobytes.

Active · paper in progress

Transfer Learning & Model Merging

Research on heterogeneous model merging and cross-architecture knowledge transfer: from merging instruction-tuned LLMs across model families to moving knowledge between fundamentally different architectures, autoregressive to diffusion.

  • #1Neo-Falcon3-7B-Instruct on the Hugging Face Open LLM Leaderboard, under-7B category
  • 54.22%Winogrande after merging Qwen2.5-0.5B weights into the Blackice diffusion language model

Neo collection

LLM merges built with heterogeneous merging techniques, combining models across different families and architectures.

Neo-Falcon3-7B-Instructagainst Falcon3-7B-Instruct
  1. MMLU42.83+1.3641.47
  2. ARC-Challenge55.18+3.0152.17
  3. Winogrande70.32+4.2666.06
  4. Hellaswag80.10+3.1077.00
Neo-Gemma-3-12B-Instructagainst gemma-3-12b-it
  1. MMLU42.76−0.8443.60
  2. ARC-Challenge51.51+1.6849.83
  3. Winogrande76.48+2.0574.43
  4. Hellaswag83.00+1.5081.50
  • the Neo merge, with its difference from the original
  • the original instruct model

Accuracy in percent, higher is better; Hellaswag on 1,000 examples. Bars start at zero and both charts share one scale, so the gains read small: the numbers carry them.

  • Neo-Falcon3-7B-Instructmerge of Falcon3-7B and Qwen2.5-7B; reached #1 on the Open LLM Leaderboard, under 7B
  • Neo-Gemma-3-12B-Instructheterogeneous merge of Gemma-3-12B
  • Neo-Mistral-8B / 3B-Instructmerges of Ministral instruction models
  • Neo-Qwen-4B-Instructmerge of Qwen3-4B

Autoregressive to diffusion

Method
Procrustes-aligned weight merging, an SVD rotation followed by per-row SLERP, from Qwen2.5-0.5B into the Blackice diffusion language model.
Result
54.22% on Winogrande, closing in on native autoregressive models at the same scale.
Finding
Feed-forward layers transfer across attention architectures: bidirectional diffusion to and from causal autoregressive.

Mistral AI Hackathon · Doc-to-LoRA · MLX

Thoth & mlx-D2L

Compresses a document's knowledge directly into model weights instead of spending context-window tokens on it. Sakana AI's Doc-to-LoRA ported to Mistral's Ministral-3B, with the full inference pipeline reimplemented in MLX for native Apple Silicon support.

Hypernetwork
Trained from scratch on Ministral-3B: 0.744 loss over 4K steps on synthetic question-answer pairs.
Thoth agent
A DSPy-based reasoning system that converts documents into LoRA adapters on the fly.
Why
Knowledge in weights keeps the context window free for reasoning, in place of traditional retrieval-augmented generation.
Runtime
A full MLX inference pipeline. No CUDA required.

Masked discrete diffusion · trained from scratch

Blackice DLM Tiny · 448M

A diffusion language model trained from scratch on a MacBook Pro M4 Max, with periodic cloud compute on Hugging Face Jobs. Absorbing-state discrete diffusion with a cosine noise schedule and bidirectional attention: generation starts from a fully masked sequence and unmasks positions in parallel, step by step, rather than left to right.

  • 448 Mparameters: 28 layers, 1,024 hidden, 16 heads
  • 24.87best ELBO perplexity, at a sequence length of 2,048
  • 2,048token context, grown 256 → 512 → 1,024 → 2,048 during training
Architecture
28 transformer layers, SwiGLU feed-forward, gated bidirectional attention, RoPE, self-conditioning, EMA; ModernBERT tokenizer.
Depth
Block attention residuals, a learned depth-wise aggregation at block boundaries; U-Net mirrored skip connections between the two halves; looped middle blocks for extra compute depth without extra parameters.
Training
FineWeb-Edu-Dedup with cosine learning-rate decay and adaptive time warping, at progressively longer context.
Also explored
LoRA and DoRA fine-tuning, SFT and GRPO, including coupled-GRPO with antithetic variates for variance reduction; a Doc-to-LoRA hypernetwork; cross-architecture weight merging; an MLX inference port; mixed-data continual pretraining on FineWeb and Cosmopedia.

Spanish legal domain · open source

Temis

A collection of language models fine-tuned on articles from the Spanish Official State Gazette (BOE), specialising in legal and administrative language. Named after Themis, the goddess of justice.

  • 8models, all trained on the BOE legal corpus
  • 357 Karticles from the Boletín Oficial del Estado
  • 360M–8Bparameters, from SmolLM to Llama 3
Base models
Gemma 2B and 7B · Llama 3.2-1B and 3-8B · Mistral 7B-v0.3 · Qwen 2.5-0.5B and 1.5B · SmolLM-360M

MLX implementation · arXiv 2602.04118

TinyLoRA

An MLX implementation of TinyLoRA (“Learning to Reason in 13 Parameters”), a parameter-efficient fine-tuning method that reaches strong task performance with extremely small adapters. Validated on Texas Hold'em poker strategy with Llama 3.2-1B: the adapter is 470 KB.

Preflop accuracyhigher is better
  1. Llama 3.2-1B, untuned7%
  2. Standard LoRA83%
  3. TinyLoRA36%
Postflop accuracyhigher is better
  1. Llama 3.2-1B, untuned13%
  2. Standard LoRA55%
  3. TinyLoRA60%
  • TinyLoRA, 470 KB adapter
  • standard LoRA fine-tuning
  • the untuned model

Accuracy on PokerBench; both charts share a 0–100% scale. Standard fine-tuning is LoRA with unsloth-mlx. TinyLoRA matches full LoRA on postflop play with orders of magnitude fewer parameters; on preflop play it recovers less.