Custom ForcingTraining-Free Subject Customization for Autoregressive Video Generation

A specific subject, kept in a streaming video for minutes, without any training.

arXiv 2026
1Kyung Hee University    2The University of Texas at Austin
*Equal contribution    †Corresponding author
Reference images
Prompt
Custom Forcing
Custom Forcing
Customized video
same subject · training-free

Abstract

Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use pretrained conditioning networks that jointly process all video frames with bidirectional attention. Neither approach is designed for causal streaming. We present Custom Forcing, a training-free method that stores reference-based anchor frames in the persistent KV cache of a frozen autoregressive video model. However, fixed anchors face two limitations: simple conditioning allows identity to drift, and the text prompt continues to favor a generic subject. To address these problems, drift-adaptive value amplification (DVA) scales reference influence with the degree of identity drift, while anchor contrast guidance (ACG) steers generation away from the generic class prior. Over two-minute rollouts, fixed anchors fall from 0.58 to 0.42 in DINO-I, while Custom Forcing keeps it between 0.58 and 0.62 without reducing motion. Custom Forcing also achieves higher subject similarity than bidirectional customization methods and better preserves identity over 30 s than causal image-to-video and reference-to-video models, while generating each frame 9.5–28.5× faster than these long-video baselines.

Top: an existing autoregressive video model generates a generic instance of the category. Bottom: Custom Forcing conditions on a few reference images and keeps the specific subject throughout the video.
Given the same prompt, an existing autoregressive video model (top) generates a generic instance of the category, whereas Custom Forcing (bottom) also conditions on a few reference images and preserves the identity of the specific subject throughout the video, without any training.

Method

Customizing a streaming model needs no new weights. The model already reads its KV cache at every step, and part of that cache, the attention sink, persists for the whole rollout. Custom Forcing writes images of the subject into this persistent part and then controls how strongly, and in which direction, the model reads them.

Overview of Custom Forcing: custom K/V construction, drift-adaptive value amplification, and anchor contrast guidance.

Custom K/V construction

The five reference images, followed by copies of one custom anchor that places the subject in the scene of the prompt, are encoded once and written into the KV cache before generation. They occupy the attention sink, so every chunk can attend to them through ordinary attention, with no retrieval step and no extra conditioning network.

Drift-adaptive value amplification

Each finalized chunk is decoded and compared with the reference images. The amount by which its similarity falls below the level of the first chunks becomes an amplification of the reference values, restricted to the subject tokens of each reference image. The reference is read more strongly as the subject drifts, instead of uniformly strongly, which would suppress motion.

Anchor contrast guidance

Inside every self-attention layer, the current chunk reads the cache twice, with and without the anchor frames. Extrapolating along the difference strengthens the anchors against the class prior, the generic instance of the category that the prompt favors. No negative prompt and no extra forward pass are needed.

Two-minute generation

With the reference images written into the cache but nothing else (Custom K/V only), the subject drifts within the first minute: the toy's eyes grow and rise on stalks, and the gray merle dog turns black tricolor. Custom Forcing keeps both subjects for the full two minutes, and its DINO-I stays between 0.58 and 0.62 while fixed anchors fall from 0.58 to 0.42.

DINO-I over two minutes: Custom Forcing stays near 0.6, Custom K/V only declines to 0.42.
DINO-I over two minutes, mean over 100 prompts.

Comparison with image- and reference-to-video models

A direct alternative gives the custom anchor to an image-to-video model as its first frame, or the reference images to a reference-to-video model. The causal models drift away from the subject within 30 s. With a 1.3B backbone, Custom Forcing matches the 14B SkyReels-V3 in whole-video DINO-I and exceeds it in the last window, while generating each frame 9.4–28.5× faster.

MethodBackboneTime / frameCausalDINO-I allfirstmiddlelastCLIP-ICLIP-T
FramePackHunyuanVideo 13B9.4×✗0.6250.6160.6200.6230.7950.357
FramePack-F1HunyuanVideo 13B9.5×✓0.4770.5340.4180.3920.7730.355
SkyReels-V2 DFWan2.1 1.3B11.5×✓0.4490.5520.3930.3070.7670.361
SkyReels-V3Wan2.1 14B28.5×✓0.6410.6930.6200.5640.7940.347
Custom Forcing (Ours)Wan2.1 1.3B1.0×✓0.6350.6190.6260.6160.8030.353
30 s videos, mean over 100 prompts. Time is the generation time per frame relative to Custom Forcing on one H100 GPU. Bold marks the best score and underline the second best.

Adding the components one at a time

With the prompt alone, the backbone generates a generic member of the category. Custom K/V only introduces the subject but drifts back toward a generic instance by 21 s. Adding DVA stops the drift and keeps DINO-I near 0.60 from the first window to the last, and adding ACG on top raises every window by a further 0.02–0.03 and keeps both subjects as in the reference images.

DVAACGDINO-I firstmiddlelast
Custom K/V only✗✗0.5790.5060.475
+ DVA✓✗0.6000.5980.598
+ DVA + ACG (Ours)✓✓0.6190.6260.616
30 s videos, mean over 100 prompts. Bold marks the best score and underline the second best.

BibTeX

@article{ok2026customforcing,
  title   = {Custom Forcing: Training-Free Subject Customization for Autoregressive Video Generation},
  author  = {Ok, Yunseung and Kim, Hyunsoo and Kim, Minseo and Kim, Suhyun},
  journal = {arXiv preprint},
  year    = {2026}
}