Error Analyses of Auto-Regressive Video Diffusion Models
Jing Wang, Fengzhuo Zhang, Xiaoli Li, Vincent Y.~ F. Tan, Tianyu Pang, Chao Du, Aixin Sun, Zhuoran Yang; 27(135):1−51, 2026.
Abstract
Auto-Regressive Video Diffusion Models (AR-VDMs) have shown strong capabilities in generating long, photorealistic videos, but suffer from two key limitations: (i) history forgetting, where the model loses track of previously generated content, and (ii) temporal degradation, where frame quality deteriorates over time. Yet a rigorous theoretical analysis of these phenomena is lacking, and existing empirical understanding remains insufficiently grounded. In this paper, we introduce Meta-ARVDM, a unified analytical framework that studies both errors through the shared autoregressive structure of AR-VDMs. We show that history forgetting is characterized by the conditional mutual information between the generated output and preceding frames, conditioned on inputs, and prove that incorporating more past frames monotonically alleviates history forgetting, thereby theoretically justifying a common belief in existing works. Moreover, our theory reveals that standard metrics fail to capture this effect, motivating a new evaluation protocol based on a “needle-in-a-haystack” task in closed-ended environments (DMLab and Minecraft). We further show that temporal degradation can be quantified by the cumulative sum of per-step errors, enabling prediction of degradation for different schedulers without video rollout. Finally, our evaluation uncovers a strong empirical correlation between history forgetting and temporal degradation, a connection not previously reported.
[abs]
[pdf][bib]| © JMLR 2026. (edit, beta) |
Facts Only
* Auto-Regressive Video Diffusion Models (AR-VDMs) show limitations in generating long, photorealistic videos.
* Two key limitations are history forgetting and temporal degradation.
* Meta-ARVDM is introduced as a unified analytical framework studying errors via AR-VDMs' autoregressive structure.
* History forgetting is characterized by conditional mutual information between output and preceding frames, conditioned on inputs.
* Incorporating more past frames monotonically alleviates history forgetting.
* Standard metrics fail to capture the effects of history forgetting and temporal degradation.
* A new evaluation protocol based on a "needle-in-a-haystack" task in DMLab and Minecraft is motivated.
* Temporal degradation can be quantified by the cumulative sum of per-step errors, allowing prediction without video rollout.
* An empirical correlation exists between history forgetting and temporal degradation.
Executive Summary
Full Take
The research establishes a fundamental connection between two distinct error modes in generative video modeling: the loss of sequential memory (history forgetting) and the decline in fidelity over time (temporal degradation). The assertion that incorporating past frames directly mitigates memory loss provides a theoretical justification for empirical improvements, shifting the understanding from mere observation to principled causality within the AR-VDM structure. The motivation to develop new evaluation protocols, like the "needle-in-a-haystack" task, suggests that existing metrics are inadequate because they treat these errors in isolation rather than acknowledging their coupled nature. The finding that these two phenomena correlate empirically challenges the practice of evaluating model components separately; one cannot analyze memory loss without considering frame quality degradation, and vice versa. This forces a reconsideration of what constitutes a meaningful metric for video generation—moving from isolated fidelity checks to holistic temporal coherence analysis.
What are the consequences if models are optimized solely against mitigating one error while ignoring the other?
How might this coupled nature fundamentally change the theoretical approach to conditioning in diffusion models beyond simple frame-by-frame prediction?
Sentinel — Human
The text exhibits the formal structure and specialized vocabulary of peer-reviewed academic writing, suggesting human authorship or direct translation/structuring of research findings.
