audio2photoreal

Avatar generation

Generating photorealistic avatars from audio

Code and dataset for photorealistic Codec Avatars driven from audio

GitHub

3k stars
31 watching
261 forks
Language: Python
last commit: about 2 years ago

Related projects:

RepositoryDescriptionStars
facebookresearch/imagebindAn AI framework that combines data from multiple sources into a single embedding space, enabling various applications such as cross-modal retrieval and generation.8,424
sebastianstarke/ai4animationA deep learning framework for data-driven character animation in Unity3D7,927
zejun-yang/aniportraitAn open-source framework for generating photorealistic animations driven by audio and reference images.4,718
facebookresearch/dinov2A PyTorch implementation of a self-supervised learning method for learning robust visual features without supervision.9,425
facebookresearch/pytorch3dA deep learning library for 3D data processing and computer vision research using PyTorch8,889
pyannote/pyannote-audioA toolkit for speaker diarization using PyTorch and speech activity detection.6,508
facebookresearch/sam2A software framework for video segmentation in images and videos using AI models13,054
facebookresearch/ca_bodyA Python implementation of a neural network architecture for image avatar body generation47
facebookresearch/audiocraftA deep learning library for generating high-quality audio21,134
huggingface/lerobotA platform providing pre-trained models, datasets, and tools for robotics with focus on imitation learning and reinforcement learning.7,874
facebookresearch/eftProvides pseudo-GT 3D human pose data and pre-trained models for training 3D pose estimation algorithms379
nvidia/vid2vidA PyTorch implementation of a video-to-video translation method for generating photorealistic videos from semantic label maps or other input data.8,623
tyiannak/pyaudioanalysisA comprehensive Python library for feature extraction, classification, segmentation, and applications of audio data.5,918
pytorch/audioA PyTorch module providing tools and functions for audio signal processing2,561
nvidia/waveglowGenerates high-quality speech from mel-spectrograms using a flow-based network architecture2,294