A specific subject, kept in a streaming video for minutes, without any training.
Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use pretrained conditioning networks that jointly process all video frames with bidirectional attention. Neither approach is designed for causal streaming. We present Custom Forcing, a training-free method that stores reference-based anchor frames in the persistent KV cache of a frozen autoregressive video model. However, fixed anchors face two limitations: simple conditioning allows identity to drift, and the text prompt continues to favor a generic subject. To address these problems, drift-adaptive value amplification (DVA) scales reference influence with the degree of identity drift, while anchor contrast guidance (ACG) steers generation away from the generic class prior. Over two-minute rollouts, fixed anchors fall from 0.58 to 0.42 in DINO-I, while Custom Forcing keeps it between 0.58 and 0.62 without reducing motion. Custom Forcing also achieves higher subject similarity than bidirectional customization methods and better preserves identity over 30 s than causal image-to-video and reference-to-video models, while generating each frame 9.5–28.5× faster than these long-video baselines.
Customizing a streaming model needs no new weights. The model already reads its KV cache at every step, and part of that cache, the attention sink, persists for the whole rollout. Custom Forcing writes images of the subject into this persistent part and then controls how strongly, and in which direction, the model reads them.
The five reference images, followed by copies of one custom anchor that places the subject in the scene of the prompt, are encoded once and written into the KV cache before generation. They occupy the attention sink, so every chunk can attend to them through ordinary attention, with no retrieval step and no extra conditioning network.
Each finalized chunk is decoded and compared with the reference images. The amount by which its similarity falls below the level of the first chunks becomes an amplification of the reference values, restricted to the subject tokens of each reference image. The reference is read more strongly as the subject drifts, instead of uniformly strongly, which would suppress motion.
Inside every self-attention layer, the current chunk reads the cache twice, with and without the anchor frames. Extrapolating along the difference strengthens the anchors against the class prior, the generic instance of the category that the prompt favors. No negative prompt and no extra forward pass are needed.
With the reference images written into the cache but nothing else (Custom K/V only), the subject drifts within the first minute: the toy's eyes grow and rise on stalks, and the gray merle dog turns black tricolor. Custom Forcing keeps both subjects for the full two minutes, and its DINO-I stays between 0.58 and 0.62 while fixed anchors fall from 0.58 to 0.42.
A direct alternative gives the custom anchor to an image-to-video model as its first frame, or the reference images to a reference-to-video model. The causal models drift away from the subject within 30 s. With a 1.3B backbone, Custom Forcing matches the 14B SkyReels-V3 in whole-video DINO-I and exceeds it in the last window, while generating each frame 9.4–28.5× faster.
| Method | Backbone | Time / frame | Causal | DINO-I all | first | middle | last | CLIP-I | CLIP-T |
|---|---|---|---|---|---|---|---|---|---|
| FramePack | HunyuanVideo 13B | 9.4× | ✗ | 0.625 | 0.616 | 0.620 | 0.623 | 0.795 | 0.357 |
| FramePack-F1 | HunyuanVideo 13B | 9.5× | ✓ | 0.477 | 0.534 | 0.418 | 0.392 | 0.773 | 0.355 |
| SkyReels-V2 DF | Wan2.1 1.3B | 11.5× | ✓ | 0.449 | 0.552 | 0.393 | 0.307 | 0.767 | 0.361 |
| SkyReels-V3 | Wan2.1 14B | 28.5× | ✓ | 0.641 | 0.693 | 0.620 | 0.564 | 0.794 | 0.347 |
| Custom Forcing (Ours) | Wan2.1 1.3B | 1.0× | ✓ | 0.635 | 0.619 | 0.626 | 0.616 | 0.803 | 0.353 |
With the prompt alone, the backbone generates a generic member of the category. Custom K/V only introduces the subject but drifts back toward a generic instance by 21 s. Adding DVA stops the drift and keeps DINO-I near 0.60 from the first window to the last, and adding ACG on top raises every window by a further 0.02–0.03 and keeps both subjects as in the reference images.
| DVA | ACG | DINO-I first | middle | last | |
|---|---|---|---|---|---|
| Custom K/V only | ✗ | ✗ | 0.579 | 0.506 | 0.475 |
| + DVA | ✓ | ✗ | 0.600 | 0.598 | 0.598 |
| + DVA + ACG (Ours) | ✓ | ✓ | 0.619 | 0.626 | 0.616 |
@article{ok2026customforcing,
title = {Custom Forcing: Training-Free Subject Customization for Autoregressive Video Generation},
author = {Ok, Yunseung and Kim, Hyunsoo and Kim, Minseo and Kim, Suhyun},
journal = {arXiv preprint},
year = {2026}
}