Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Tune-A-Video enables one-shot tuning of image diffusion models to generate text-to-video, by fine-tuning on a single text-video pair rather than large video datasets. It builds on pretrained text-to-image diffusion models and introduces a tailored spatio-temporal attention mechanism with an efficient one-shot tuning strategy, using DDIM inversion at inference for structure guidance to learn coherent motion. The approach supports editing prompts to produce stylized videos and provides code, demos, and a published model for practical use.
Parse Score