Technical paper
Coming soon.
A paper with the full method and measurements behind the Niwaki models is in preparation.
It will cover how the models are pruned and recovered, the evaluation protocol — perplexity, task suites, and the program-state behaviours that compression tends to break first — and how pruning compares with quantizing the unpruned model at the same size.
Until it is out, the model cards are the record: each one states its protocol and carries the full tables for its release.