The baseline JiT predicts clean pixels (x_0) directly. REPA adds an intermediate DINO patch-alignment loss, but its projected features are not passed to later denoising blocks. Standard REG jointly denoises a DINO token, yet predicts its clean state only after the final block.
New Experiment instead jointly diffuses 17 DINOv3 tokens (CLS + pooled 4×4 spatial tokens) and
predicts their clean semantic (x_0) after block 4. The clean CLS prediction conditions the
remaining blocks, while a wide CLS latent is compressed into four continuation tokens. This