Engraving of a Japanese painter’s low worktable with an unfurled pine-and-mountain landscape, brushes, an inkstone and bowls of mineral pigment, under the word Nihonga.

Research notebook · updated

Nihonga 日本画

A smaller image model, with every cut measured.

We are compressing Qwen-Image-2.1: identify transformer blocks that can be replaced, learn inexpensive bridges, and distill the original model’s behavior back into the student. This is our working research record, including results that did not work.

Phases 0–2 · establish the reference

Understand the model before removing anything

The inspected transformer has 32 blocks, a hidden width of 4,096, and approximately 7.115 billion parameters. The 60-layer examples in the original plan are illustrative; our experiments use this actual 32-block architecture. The model predicts a velocity that updates the noisy latent at each denoising step; the VAE decodes the final latent into an image.

What “internal” means

Internal means our own synthetic prompt dataset, built for this compression project. It is separate from Qwen-Image-Bench. The “Internal” baseline row is specifically its 1,007-prompt held-out internal-test split, not the training set and not all 20,000 pilot prompts. Qwen-Image-Bench supplies the other two rows, with the same 1,000 benchmark items evaluated in Chinese and English.

Why build our own dataset?

We need prompts to train the bridges and heal the student without training on benchmark questions. We also need validation prompts for choosing checkpoints and a separate test set for evaluating changes. A broad synthetic corpus lets us deliberately cover counting, spatial relations, text rendering, materials, scenes and visual styles instead of relying on whichever examples happen to be easy. For distillation, the original image model supplies hidden states and velocity targets; real-image training pairs are not required for this stage.

How we created it

1 · structured prompts
Generate a 100,000-prompt template pool across five capability groups and 38 subdimensions. Combine objects, counts, colors, relations, materials and scenes with difficulty levels 1–5, styles, prompt formats and aspect ratios. Exact normalized duplicate prompts are removed. Counting and spatial relations receive extra weight, and text rendering receives twice its unadjusted subdimension weight.
2 · stable partitions
Assign deterministic IDs and seeds, stratify within capability/subdimension/difficulty, then select a 20,000-prompt prefix by a saved pool rank. The intended split is approximately 90% train, 5% validation and 5% internal test. The actual pilot sizes are 17,991 / 1,002 / 1,007. Split assignments remain fixed when growing the pool.
3 · natural language
Use Qwen3-8B through vLLM on Modal to rewrite selected English prompts and translate about 10% into Chinese. Pure templates can cover a narrow set of sentence shapes; natural wording and Chinese prompts broaden the conditioning the student sees. Automatic checks try to preserve quoted text and key constraints; failed rewrites fall back to the original template. Raw responses are retained.
4 · benchmark separation
Check normalized exact matches and English 6-word / Chinese 8-character n-gram overlap against a pinned Qwen-Image-Bench prompt snapshot, both before and after rewriting. At the configured 0.5 overlap threshold, the template-pool check removed zero prompts; no final rewrites were rejected for benchmark overlap. These checks reduce detectable contamination, but do not prove every prompt is semantically unrelated.
5 · serialize the record
Save each prompt’s capability, subdimension, facet, difficulty, style, aspect ratio, language, source, original template, ID, seed and split. Save the generation configuration, raw LLM outputs, overlap checks, distributions and artifact hashes. Rerunning reads the existing results rather than generating the corpus again.
Our dataset · split sizesnumber of prompts · 20,000 total
  1. Train · fit parameters17991
  2. Validation · select checkpoints1002
  3. Internal test · held-out evaluation1007
Pilot prompt capabilitiesshare of 20,000 prompts
  1. Alignment25.980%
  2. Creative20.000%
  3. Aesthetics18.020%
  4. Visual Quality18.010%
  5. Real World17.990%
How the final pilot is writtenshare of final pilot prompts
  1. Retained template wording60.910%
  2. LLM English rewrite29.070%
  3. LLM Chinese translation10.020%
Pilot languagesshare of final pilot prompts
  1. English89.980%
  2. Chinese10.020%

The 100,000-prompt master pool is the template reservoir. Natural-language rewriting and translation were completed for the 20,000-prompt pilot; we have not trained on all 100,000 prompts. The bars show the actual pilot distributions from the saved Phase 1 statistics, not planned quotas. Wording checks are automatic rather than a guarantee of perfect constraint preservation.

A saved internal-test example

Create a photorealistic café chalkboard displaying “Today’s special: tomato soup and grilled cheese” in legible chalk writing, soft shadows, mysterious atmosphere.

This held-out example tests long text rendering: creative generation, difficulty 5, photographic style, square aspect ratio, English LLM rewrite. Its seed and original template are saved. It is a test prompt, not a bridge-training example.

“Internal” in the baseline is the held-out test row; validation is a different split
Split / benchmarkRoleHow used so far
Train · 17,991 promptsFit model parameters20 selected prompts in the bridge and healing pilots
Validation · 1,002 promptsChoose candidates and checkpoints10 selected prompts, reused across the small pilots
Internal test · 1,007 promptsOur held-out evaluation setAll baseline images; subsets for Phase 3 diagnostics
Qwen-Image-Bench · 1,000 bilingual itemsExternal benchmark · evaluation only1,000 Chinese and 1,000 English baseline images

The original image model baseline

We generated and judged 3,007 reference images: 1,000 Chinese benchmark prompts, 1,000 English benchmark prompts and 1,007 internal prompts. Another 301 paired regenerations measure generation-and-judge variation. Prompts, seeds, resolutions, scheduler settings, model revisions and judge outputs are saved.

Original model · overall judge score0–100 · higher is better
  1. Bench Cn53.615
  2. Bench En53.669
  3. Internal56.084
Original model scores · frozen local Q-Judger protocol · 0–100
Evaluation setImagesOverallPrompt adherenceText rendering
Bench Cn100053.6254.6443.59
Bench En100053.6753.8841.69
Internal100756.0857.1370.49

These are our local judge scores, not a claim of an official leaderboard result. Text scores cover only applicable prompts; contributing counts are included in the downloadable data. Judge calibration found an approximately 2.5-point aesthetics offset; the judge remains frozen for paired comparisons.

Baseline speed
6.172 seconds per image on an H100 80 GB, batch size 1, 40 denoising steps; 80 timed calls across eight aspect ratios after warm-up.
Baseline memory
36.92 GiB peak allocated VRAM across those calls. Student inference has not yet been benchmarked under the same protocol.

Phase 3 · measure the missing update

Similarity is a clue; images decide what it misses

We measured all 32 transformer blocks before choosing which ones to replace. On 64 prompts, at three denoising stages, we compared the hidden states entering and leaving every block. Blocks 2, 3, 4 and 5 had the smallest consistent changes, so we chose them as bridge candidates. The choice came from the measurements.

To make that choice, we took each block’s average relative L2 change at each of the three stages, kept the largest of those three averages, and ranked all 32 blocks from smallest to largest. The first four were blocks 2, 4, 5 and 3. After selecting them, we temporarily skipped each of those four blocks separately to measure how much that changed the model’s velocity prediction. The all-block activation comparison and the four-candidate bypass tests are two different experiments.

The initial candidate ranking

Initial shortlist · rank all 32 blocks by max(early mean, middle mean, final mean) · lower is smaller
RankOriginal blockLargest stage-average L2 changeDecision
128.681%Initial candidate
249.574%Initial candidate
3510.043%Initial candidate
4310.881%Initial candidate
5111.740%Comparison context
6615.500%Comparison context
7916.504%Comparison context
81016.912%Comparison context

This rule favors blocks with consistently small updates across the sampled denoising stages, rather than a block that appears unimportant at only one stage. It is a ranking of local hidden-state changes, not yet bypass velocity errors or image quality. “Largest stage average” means the largest of three means over 64 prompts; it is not the worst individual prompt. The initial rank order is 2, 4, 5, 3; written in model order, the selected set is 2–5.

The activation measurements

We measured all 32 blocks on 64 fixed internal prompts at denoising steps 0, 19 and 39 of the 40-step teacher: 6,144 block/stage observations. Hooks compare image-token states immediately before and after each block, excluding the text prefix. Each cosine is averaged over image tokens, then across prompts. The table shows the four tested candidates; the graph and downloadable record cover every block.

Direction · cosine, zoomedcloser to 1 means less rotation
Cosine for blocks 1–6, with a zoomed vertical axis from 0.994 to 1.000. Blocks 2–5 are highlighted. Open the SVG for a larger view.
Update size · relative L2smaller means a smaller update
Relative L2 hidden-state change for blocks 1–6, on a linear vertical axis from 0% to 18%. Blocks 2–5 are highlighted. Open the SVG for a larger view.
  • early · dashed
  • middle · dotted
  • final · solid

These views zoom into blocks 1–6, including a neighbor on each side of the tested 2–5 batch. The cosine axis is deliberately narrowed to 0.994–1.000; its larger-looking differences are still small absolute changes. Relative L2 is ‖h_out − h_in‖₂ / ‖h_in‖₂, reported as a percentage on a separate linear axis. Cosine compares direction; L2 also responds to changes in magnitude. The green wash identifies the tested batch, not an acceptance region.

Show all 32 blocks · full-range cosine and logarithmic relative L2
Cosine · full rangevertical axis 0–1
Full-range cosine across all 32 blocks at three stages, retaining the large changes at the first and final blocks.
Relative L2 · all blockslogarithmic axis · 5%–1,500%
Relative L2 changes across all 32 blocks, on a logarithmic axis from 5% to 1,500%, so the large block-0 update does not flatten the remaining blocks.

The L2 overview uses a log scale: equal vertical distances represent equal ratios rather than equal percentage-point changes. It preserves block 0’s update, which exceeds 800% of its input norm at the early stage, while keeping other blocks visible. All plotted observations fit the stated axes; none are clipped. The stage legend above applies to both overviews.

Mean input/output image-token cosine · 1 means same direction
Original blockEarly cosineMiddle cosineFinal cosine
20.9983210.9985970.997599
30.9971100.9981590.997847
40.9969890.9972810.997353
50.9958300.9961610.996551
Relative hidden-state change · ‖h_out − h_in‖₂ / ‖h_in‖₂ · lower is smaller
Original blockEarly changeMiddle changeFinal change
28.681%5.544%8.126%
310.881%7.473%8.003%
48.460%9.322%9.574%
59.916%10.043%10.034%

For example, block 2’s early cosine is 0.998321, but the update still has a norm equal to 8.681% of its input’s norm. High cosine does not imply a negligible update or 99.8% image quality. Block 1 and some later blocks also have high cosine; the shortlist was determined by the maximum stage-average relative L2 change, rather than cosine alone. These four earn a first test, not automatic approval for removal. Prompt-level distributions and per-capability aggregates are in the measurements.

What happens when each update is removed?

An identity bypass returns h_out = h_in, so the missing block contributes no update. We tested each candidate separately on the same five pilot prompts × three stages, using saved teacher-trajectory inputs and rebuilt student prefix caches. Each block has 15 velocity comparisons, with intact-teacher replay controls matching exactly within the diagnostic. These five-prompt Phase 3 probes are distinct from the later ten-prompt bridge-validation set.

Early stage · identity bypasslower is better
  1. Skip block 216.407%
  2. Skip block 315.801%
  3. Skip block 410.811%
  4. Skip block 512.740%
Middle stage · identity bypasslower is better
  1. Skip block 24.145%
  2. Skip block 33.442%
  3. Skip block 42.687%
  4. Skip block 52.813%
Final stage · identity bypasslower is better
  1. Skip block 218.962%
  2. Skip block 311.854%
  3. Skip block 45.433%
  4. Skip block 53.687%
Phase 3 identity-bypass velocity error · five prompts · lower is better
Skipped blockEarly errorMiddle errorFinal errorWorst case
216.407%4.145%18.962%21.396%
315.801%3.442%11.854%18.531%
410.811%2.687%5.433%17.797%
512.740%2.813%3.687%17.003%

Block 2 has the strongest activation similarity of the tested candidates early and in the middle, yet its omission causes the largest mean velocity errors at all three stages. Block 4 has the lowest mean bypass error early and in the middle; block 5 is lowest at the final stage. Similarity helps nominate experiments, but bypass sensitivity and finished images change the ranking.

The worst-case column is the largest error among that block’s 15 prompt/stage cases. Velocity error is a prediction difference, not a percentage of image-quality loss. These individual interventions do not establish that several blocks can be removed together, and no bridge is fitted in this diagnostic.

What the first image comparisons showed

Five prompts, one per capability, were generated with block 4 or block 5 skipped: ten bypass images. In the astronaut example, skipping block 4 loses the spacesuit. Skipping block 5 retains it. That initially favors block 5 for closer study, while all four candidates remain eligible for learned bridges.

Each sheet reads teacher / skip block 4 / skip block 5, left to right. These are identity-bypass images from Phase 3, before bridge training or healing. Teacher images came from an earlier run; unresolved cross-run variation limits causal pixel comparisons. Observations are manual, not a blinded quality study. Open any sheet at full resolution.

Pixel MAE and PSNR record changes in composition and appearance, but do not directly measure prompt adherence or aesthetic quality. None of these identity bypasses is accepted as the finished model.

Phases 4–5 · recover the missing update cheaply

A small residual bridge in place of a full block

The bridge projects 4,096 features down to 256, applies GELU, projects back up, and adds the input: h′ = h + W_up GELU(W_down h). Each bridge has 2,097,152 parameters. We fit four independent candidates, replacing one original block at a time; the four bridges are not inserted together.

Before · the original transformer block
Original block: 4,096-feature input, normalization, attention with a gated residual, normalization, feed-forward network with a gated residual, then a 4,096-feature output.
After · a residual bottleneck bridge
Replacement: preserve the input on a residual path; project 4,096 features down to 256, apply GELU, project back to a 4,096-feature correction, then add the original input.

The student keeps the same 4,096-feature input and output interface. Only the correction travels through the 256-feature bottleneck; the residual path keeps the original information. GELU is a nonlinear transformation of those 256 features. The bridge processes each token separately without attention and learns to approximate the removed block’s update. The diagrams show shapes and operations, not identical teacher and student values or measured speedups. Open either diagram for the standalone SVG.

Bridge-only pretraining keeps the rest of the model frozen. Each candidate gets 500 updates using cached hidden states from 20 training prompts, with 10 validation prompts selecting checkpoints. The fitting objective balances prompt tokens and early, middle and late image tokens, combining normalized direction matching with a smaller relative squared-error term. Block 5’s selected checkpoint is update 300.

Does a better hidden-state fit improve the whole prediction?

We replay each identity and trained student against the verified teacher, using the same saved inputs at three stages. Each candidate has 30 comparisons; there are 240 student forwards across four candidates and two variants.

Block 2 replacementlower is better
  1. Identity bypass12.879%
  2. Pretrained bridge12.023%
Block 3 replacementlower is better
  1. Identity bypass10.637%
  2. Pretrained bridge10.238%
Block 4 replacementlower is better
  1. Identity bypass7.160%
  2. Pretrained bridge5.873%
Block 5 replacementlower is better
  1. Identity bypass6.406%
  2. Pretrained bridge5.518%
  • identity bypass
  • pretrained bridge
Mean validation velocity error · lower is better
Original blockIdentity errorBridge errorRelative reductionCases improved
212.879%12.023%6.64%27 / 30
310.637%10.238%3.74%20 / 30
47.160%5.873%17.97%26 / 30
56.406%5.518%13.86%24 / 30

How to read the percentages. Velocity error is ‖student − teacher‖₂ / ‖teacher‖₂, averaged equally across prompts and stages. For block 2, 12.8788036% becomes 12.0234947%: a drop of 0.855309 percentage points. Dividing that drop by the original 12.8788036% gives a 6.64% relative reduction. This is a fraction of prediction error removed, not an image-quality improvement. “27 / 30” counts the cases with lower error.

Block 5 has the lowest mean trained error, 5.518%, and a worst case of 15.095%, so it is selected for the healing pilot. Block 4 removes a larger fraction of its identity error but ends at a slightly higher mean, 5.873%. Block 3 slightly regresses at the early and middle stages despite improving its overall mean. The small reused validation set does not establish statistical superiority.

Phase 6 · first make the supervision repeatable

A replay discrepancy paused training

Fresh teacher predictions initially disagreed with historical saved velocities beyond our 0.1% relative-L2 tolerance. Repeating a forward in the same process was stable, and full tensor-state fingerprints matched across fresh processes. The exact cause of the historical discrepancy remains unresolved.

The numerical contract we verified

We pinned the model revision and transformer source, H100 hardware and software versions, BF16 weights, deterministic algorithms, math attention, TF32 and reduction settings, exact input order and cache lifecycle. Each prompt gets a fresh cache; early-stage conditioning is extracted before cached middle and late forwards. Reference files are loaded after predictions are made. The audit covers 301 state tensors, including otherwise unregistered positional tables.

Under this contract, all 90 teacher predictions match the new saved references exactly. Historical targets and weights remain preserved. This establishes a working repeatable setup; it does not isolate which individual setting caused the old mismatch or prove that every setting is necessary.

Did the bridges need to be trained again?

We captured 360 fresh block/stage hidden-state pairs for the same 30 prompts and token selections. Frozen bridge validation objectives differ from the historical objectives by less than 0.05%. The replay change alone therefore does not justify replacing the Phase 5 training results.

Fresh hidden-state fit check · this is a different metric from velocity error
BlockFresh hidden-fit objectiveReduction vs identityChange vs historical fit
20.00396945.53%-0.0082%
30.00521338.00%-0.0432%
40.00503749.70%+0.0175%
50.00671261.40%-0.0273%

The hidden-fit objective combines direction and relative squared errors; it has no direct percentage interpretation. Its 38–61% reductions cannot be compared as if they were the same metric as the 3.74–17.97% velocity-error reductions.

Phase 6 · two completed parameter-efficient trials

Training works; this recipe does not improve validation

Healing is a second training stage, after bridge pretraining. Phase 5 trains only the bridge to reproduce the removed block’s hidden-state update. Phase 6 starts from that pretrained bridge and trains the modified student to reproduce the teacher’s final velocity prediction, with hidden-state supervision as an auxiliary loss. “Before healing” already includes the trained Phase 5 bridge.

In these healing pilots, we trained the selected bridge together with rank-8 LoRA adapters on the query, key, value and output projections of all 31 surviving attention blocks. That is 124 adapters and 10,223,616 trainable parameters. Surviving base weights stay frozen. This is a parameter-efficient implementation pilot; full-weight architecture healing is still incomplete.

What is LoRA?

LoRA means low-rank adaptation. Instead of updating a large existing weight matrix W, we freeze it and learn a small correction through two much smaller matrices. A projection becomes y = W x + (α/r) B A x. The original computation stays; the added path learns how to adjust it. Here the rank r and scaling parameter α are both 8, so the scale is 1.

For one 4,096 → 4,096 attention projection, the original matrix has 16,777,216 weights. The rank-8 correction uses 4,096 → 8 → 4,096, with 65,536 trainable weights: 256 times fewer trainable weights than updating that whole matrix. Unlike the rank-256 bridge, the LoRA correction has no GELU. It adjusts a surviving projection; it does not replace the missing transformer block.

We add four such corrections per surviving block: query, key, value and attention output. Across 31 blocks, that is 124 LoRA modules and 8,126,464 LoRA parameters. Adding the bridge’s 2,097,152 gives 10,223,616 trainable parameters. The up-projection of every LoRA starts at zero, so the initial correction is zero and the student initially behaves exactly like the pretrained-bridge student.

How this healing differs from training only the bridge

Different training scopes and targets · bridge-only healing remains an untested ablation
Training stageWhat can changeWhat we ask it to match
Phase 5 · bridge pretrainingBridge only · 2,097,152 parametersHidden state after the removed block
Phase 6 · this healing pilotBridge + LoRA · 10,223,616 parametersWhole-model velocity + auxiliary hidden states
Bridge-only healing · not runBridge only; original surviving weights frozenWhole-model velocity + the same auxiliary losses

Phase 5 teaches the replacement to approximate a local block update. Healing gives the rest of the modified model a way to adjust to that imperfect replacement, while training against the teacher’s whole-model velocity at each sampled denoising step. The original surviving weights are frozen, but their attention computations can change through LoRA. The teacher remains frozen throughout; a velocity prediction is not a finished image.

We have not run an otherwise-identical Phase 6 experiment with only the bridge trainable. Consequently, the observed regression does not isolate LoRA as the cause. It shows that these two bridge-plus-LoRA healing recipes fail to improve this validation set; bridge-only output healing is a separate comparison still to be tested.

Bridge pretraining
Trainable: the bridge alone. Target: the hidden states after the original removed block. Purpose: approximate that block’s local update.
Healing pilot
Trainable: the pretrained bridge plus LoRA adapters in surviving blocks. Target: teacher velocity output, with auxiliary hidden-state and bridge losses. Purpose: help the whole modified model recover.
What failed
The second stage increased held-out velocity error. This is a comparison of the same student before and after extra training, not two interchangeable names for bridge pretraining. “Healing” names the intended recovery procedure; it does not guarantee that recovery occurs.
Training
120 updates per run: two shuffled passes over 20 training prompts × three stages. Same initialization and sampling seeds, data, losses and validation rule; learning rates 0.0001 and 0.00001.
Supervision
Relative velocity MSE + 0.1 normalized hidden-state loss at selected surviving blocks + 0.1 bridge-output supervision. Validation prompts never contribute optimizer gradients.
Cache gradients
Early conditioning is rebuilt with current student parameters. Checkpoint recomputation uses private caches; cached-stage gradients reach prefix keys and values without changing the shared cache during backward.
Controls
Both fresh runs passed 90 exact teacher replays, reproduced all 30 baseline student predictions, passed cached-stage gradient checks, and verified 288 surviving base tensors unchanged. Checkpoints include optimizer and RNG state; results are reused on reruns.
Healing validation trajectorylower is better
Both learning rates remain above the pretrained bridge’s 5.518% validation error across 120 updates; exact values follow.
  • 0.0001 learning rate · dashed
  • 0.00001 learning rate · solid
  • pretrained baseline · dotted
Mean validation velocity error · same 10 prompts × three stages
Optimizer updatesLearning rate 0.0001Learning rate 0.00001
05.518%5.518%
308.265%6.042%
608.781%6.186%
9010.283%6.107%
1207.188%6.113%

Decision: keep update 0 in both trials. The pretrained bridge remains at 5.518%. At update 120, the first run reaches 7.188% and the smaller-rate run 6.113%; both are worse. Lowering the rate reduces the damage, but does not provide a successful healed checkpoint. The result is preserved rather than promoted as an improvement.

A “passed” run status means the execution and integrity controls passed, not that quality improved. Two rates on 20 training prompts do not isolate the reason for regression. Objective scaling, optimization and generalization still need investigation. No finished images, independent-test scores, or inference speedups have been measured for the bridge or healed student.

The research record

What is complete, what remains open

Phase 0
Architecture inspected: 32 blocks and 4,096-wide hidden states.
Phase 1
20k prompt pilot serialized, with a 100k template pool and fixed split assignments.
Phase 2
Original-model images, judged baseline, repeatability analysis and H100 efficiency measurements saved.
Phase 3
Layer diagnostics and identity-bypass image pilots; blocks 2–5 retained as learned-replacement candidates.
Phase 4
Residual bottleneck bridges implemented and insertion checked.
Phase 5
Four independent bridge-pretraining pilots completed; checkpoints and optimizer states saved.
Phase 6
Repeatable teacher contract verified; fixed-input bridge comparisons and fresh hidden targets saved. Two healing pilots completed without validation improvement. Full architecture healing remains open.
Later phases
Broader healing coverage, finished-image benchmarking, progressive pruning, denoising-step distillation and optional quantization remain planned. No results are claimed for them.

Reproducibility and updates

Each executable notebook step has a method description and a plain-language explanation of its result. Targets, raw outputs, metrics, selected weights, optimizer state and manifests are serialized locally, with expensive experiment artifacts backed up remotely. Matching saved steps reuse their results; changed inputs stop rather than silently overwrite the record.

This page is a dated snapshot of saved experiments, not a live training dashboard. Its refresh script reads selected local summaries and copies existing comparison images; it runs no models and publishes no weights, caches or credentials. Every plotted value is available in the downloadable record.

Current result · development pilot

The bridge helps. Healing has not helped yet.

Replacing original block 5 with a pretrained bottleneck bridge reduces the student’s mean velocity error from 6.406% to 5.518% against the original model. Both 120-update healing trials made validation worse, so we retain the pretrained bridge. Image quality and speed of this student still need to be measured.

  • 31 + 1surviving transformer blocks plus one bridge, from 32 original blocks
  • 5.518%mean fixed-input validation velocity error for the selected bridge
  • 13.86%relative error reduction against skipping block 5
  • 90 / 90teacher replay predictions exactly match the verified references

The validation set is 10 prompts at three denoising stages: 30 comparisons, not 30 independent prompts. It has also been used for checkpoint and candidate selection. These numbers measure predictions on saved teacher inputs, rather than completed images or independent test generalization.