StyleVideo: One-shot Image-guided Video Style Transfer
-
Abstract
Existing video style transfer methods mainly focus on feature fusion using models trained on artistic datasets, yielding suboptimal results when provided with arbitrary style references as input. Drawing inspiration from the remarkable performance of pre-trained text-to-image diffusion models, integrating these large-scale pre-trained models into video style transfer holds the potential to create artworks with a more pronounced and discernible artistic essence. However, many diffusion-based video editing methods rely on text prompts, which presents a challenge in effectively conveying the specific artistic style inherent in an image. To address this issue, this paper presents StyleVideo, a novel one-shot image-guided video style transfer method that enables the manipulation of video styles using a single reference image. Our framework acquires a style-aware embedding from the style image through a one-shot inversion process, effectively disentangling the style information with a crafted style loss function. A conditional video diffusion model is introduced to achieve video-to-video style injection, optimizing frame fusion attention to maintain temporal coherence within the source videos. Additionally, a frame interpolation module is further incorporated to mitigate flickering issues during the style injection process, leading to smoother outcomes. Extensive experiments on a diverse set of references, covering various styles, have demonstrated the effectiveness of our method.
Full Text Translation (powered by iFLYTEK)
Full Text Translation (powered by iFLYTEK)
-
-