Dojo
A dōjō (道場) is the place of the way: the hall where a craft is practised until it becomes one's own.
Niwaki is one long project. This page keeps the others from the same bench: model merging, documents folded into weights, a diffusion language model trained from scratch, Spanish legal models, and adapters of a few hundred kilobytes.
Active · paper in progress
Transfer Learning & Model Merging
Research on heterogeneous model merging and cross-architecture knowledge transfer: from merging instruction-tuned LLMs across model families to moving knowledge between fundamentally different architectures, autoregressive to diffusion.
- #1Neo-Falcon3-7B-Instruct on the Hugging Face Open LLM Leaderboard, under-7B category
- 54.22%Winogrande after merging Qwen2.5-0.5B weights into the Blackice diffusion language model
Neo collection
LLM merges built with heterogeneous merging techniques, combining models across different families and architectures.
- the Neo merge, with its difference from the original
- the original instruct model
Accuracy in percent, higher is better; Hellaswag on 1,000 examples. Bars start at zero and both charts share one scale, so the gains read small: the numbers carry them.
- Neo-Falcon3-7B-Instructmerge of Falcon3-7B and Qwen2.5-7B; reached #1 on the Open LLM Leaderboard, under 7B
- Neo-Gemma-3-12B-Instructheterogeneous merge of Gemma-3-12B
- Neo-Mistral-8B / 3B-Instructmerges of Ministral instruction models
- Neo-Qwen-4B-Instructmerge of Qwen3-4B
Autoregressive to diffusion
- Method
- Procrustes-aligned weight merging, an SVD rotation followed by per-row SLERP, from Qwen2.5-0.5B into the Blackice diffusion language model.
- Result
- 54.22% on Winogrande, closing in on native autoregressive models at the same scale.
- Finding
- Feed-forward layers transfer across attention architectures: bidirectional diffusion to and from causal autoregressive.
Mistral AI Hackathon · Doc-to-LoRA · MLX
Thoth & mlx-D2L
Compresses a document's knowledge directly into model weights instead of spending context-window tokens on it. Sakana AI's Doc-to-LoRA ported to Mistral's Ministral-3B, with the full inference pipeline reimplemented in MLX for native Apple Silicon support.
- Hypernetwork
- Trained from scratch on Ministral-3B: 0.744 loss over 4K steps on synthetic question-answer pairs.
- Thoth agent
- A DSPy-based reasoning system that converts documents into LoRA adapters on the fly.
- Why
- Knowledge in weights keeps the context window free for reasoning, in place of traditional retrieval-augmented generation.
- Runtime
- A full MLX inference pipeline. No CUDA required.
Masked discrete diffusion · trained from scratch
Blackice DLM Tiny · 448M
A diffusion language model trained from scratch on a MacBook Pro M4 Max, with periodic cloud compute on Hugging Face Jobs. Absorbing-state discrete diffusion with a cosine noise schedule and bidirectional attention: generation starts from a fully masked sequence and unmasks positions in parallel, step by step, rather than left to right.
- 448 Mparameters: 28 layers, 1,024 hidden, 16 heads
- 24.87best ELBO perplexity, at a sequence length of 2,048
- 2,048token context, grown 256 → 512 → 1,024 → 2,048 during training
- Architecture
- 28 transformer layers, SwiGLU feed-forward, gated bidirectional attention, RoPE, self-conditioning, EMA; ModernBERT tokenizer.
- Depth
- Block attention residuals, a learned depth-wise aggregation at block boundaries; U-Net mirrored skip connections between the two halves; looped middle blocks for extra compute depth without extra parameters.
- Training
- FineWeb-Edu-Dedup with cosine learning-rate decay and adaptive time warping, at progressively longer context.
- Also explored
- LoRA and DoRA fine-tuning, SFT and GRPO, including coupled-GRPO with antithetic variates for variance reduction; a Doc-to-LoRA hypernetwork; cross-architecture weight merging; an MLX inference port; mixed-data continual pretraining on FineWeb and Cosmopedia.
Spanish legal domain · open source
Temis
A collection of language models fine-tuned on articles from the Spanish Official State Gazette (BOE), specialising in legal and administrative language. Named after Themis, the goddess of justice.
- 8models, all trained on the BOE legal corpus
- 357 Karticles from the Boletín Oficial del Estado
- 360M–8Bparameters, from SmolLM to Llama 3
- Base models
- Gemma 2B and 7B · Llama 3.2-1B and 3-8B · Mistral 7B-v0.3 · Qwen 2.5-0.5B and 1.5B · SmolLM-360M
MLX implementation · arXiv 2602.04118
TinyLoRA
An MLX implementation of TinyLoRA (“Learning to Reason in 13 Parameters”), a parameter-efficient fine-tuning method that reaches strong task performance with extremely small adapters. Validated on Texas Hold'em poker strategy with Llama 3.2-1B: the adapter is 470 KB.
- TinyLoRA, 470 KB adapter
- standard LoRA fine-tuning
- the untuned model
Accuracy on PokerBench; both charts share a 0–100% scale. Standard fine-tuning is LoRA with unsloth-mlx. TinyLoRA matches full LoRA on postflop play with orders of magnitude fewer parameters; on preflop play it recovers less.