Co-occurring keywords
Papers
SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language Models
CVPR 2024
CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation
CVPR 2024
Template Free Reconstruction of Human-object Interaction with Procedural Interaction Generation
CVPR 2024
Residual Denoising Diffusion Models
CVPR 2024
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
AAAI 2024
Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model
EMNLP 2024
Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training
EMNLP 2024
Generative Powers of Ten
CVPR 2024