Data as of Jul 25, 2026 · Based on 2,718,867 AI responses across 9,511 prompts · See how Parse measures this
Vision XL is a diffusion-model-based framework that solves high-definition video inverse problems using latent diffusion models (SDXL), enabling tasks such as deblurring, super-resolution, and inpainting without relying on external pre-trained modules. It introduces pseudo-batch consistent sampling and pseudo-batch inversion to improve temporal consistency and GPU efficiency, allowing 25-frame videos to be processed on a single GPU with about 13 GB VRAM in roughly 2.5 minutes. The approach achieves state-of-the-art video reconstruction across a wide range of spatio-temporal degradations and is built on open-source latent diffusion models.
Parse Score