This survey provides a comprehensive overview of controllable video generation, a rapidly developing subfield of AI-generated content (AIGC). As AI models become more sophisticated, there's a growing need for methods that allow users to precisely control video output to match their intent. Current text-to-video models often fall short because text prompts alone are insufficient for complex or fine-grained requests. To address this, researchers are integrating non-textual conditions, such as camera motion or human pose, into existing video generation models to enable more accurate and flexible video synthesis. The survey systematically reviews the theoretical foundations and recent advancements in controllable video generation. It covers key concepts, common open-source video generation models, and delves into control mechanisms within diffusion models. Methods are categorized based on the types of control signals used, including single-condition, multi-condition, and universal controllable generation. The goal is to offer a clear understanding of the field's progress and future directions.
This survey covers the theoretical foundations and recent advances in video generation models relevant to controllable video generation.
Controllable video generation has significant implications for various creative industries and research roles, enhancing the practical applicability of AI in content creation.
This document is a survey paper and does not contain information on fees.
The survey focuses on enhancing user control in AI video generation by integrating non-textual conditions into existing video generation models. It reviews the theoretical foundations and recent advancements in this subfield of AI-generated content.
Graduates could pursue roles such as AI Researcher, Machine Learning Engineer, Computer Vision Engineer, Content Creator, Video Editor, and Generative AI Specialist. This field enhances the practical applicability of AI in various creative industries and research roles.
The survey covers key concepts and theoretical foundations of various video generative models, including GANs, VAEs, Flows, Diffusion Models, and Autoregressive Models.