GPT-SoVITS

Voice Generator

An AI system for generating human-like voices from text inputs, using deep learning techniques and pre-trained models.

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

GitHub

37k stars
218 watching
4k forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list

text-to-speechttsvitsvoice-clonevoice-cloneaivoice-cloning

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
jasonppy/voicecraftA neural codec model for speech editing and text-to-speech synthesis in real-time, using few seconds of reference audio.7,744
metavoiceio/metavoice-srcA deep learning model for generating human-like speech3,936
neonbjb/tortoise-ttsAn open-source text-to-speech system trained with high-quality audio capabilities13,373
coqui-ai/ttsA deep learning toolkit for generating human-like speech from text36,118
plachtaa/vall-e-xA research implementation of Microsoft's VALL-E X zero-shot TTS model for multilingual text-to-speech synthesis and voice cloning7,719
coqui-ai/sttA toolkit for building and deploying speech-to-text models using deep learning techniques2,302
mozilla/ttsAn open-source project providing a suite of deep learning models and tools for advanced text-to-speech synthesis.9,466
mshumer/gpt-prompt-engineerA tool for automating the process of generating and ranking effective prompts for AI models like GPT-4, GPT-3.5-Turbo, or Claude 3 Opus.9,411
tensorspeech/tensorflowttsReal-time speech synthesis using state-of-the-art architectures3,855
openai/whisperA general-purpose speech recognition system trained on large-scale weak supervision72,752
instruction-tuning-with-gpt-4/gpt-4-llmThis project generates instruction-following data using GPT-4 to fine-tune large language models for real-world tasks.4,244
jaywalnut310/vitsDevelops an end-to-end text-to-speech system that generates more natural audio than existing models6,947
minimaxir/gpt-2-simpleA tool for retraining and fine-tuning the OpenAI GPT-2 text generation model on new datasets.3,398
camb-ai/mars5-ttsA deep learning-based speech synthesis model that generates natural-sounding audio with controlled prosody.2,551
rhasspy/piperA fast local neural text-to-speech system optimized for small devices7,002