VoiceCraft
by jasonppy
Zero-Shot Speech Editing and Text-to-Speech in the Wild
AI summary
Speech editor
A neural codec model for speech editing and text-to-speech synthesis in real-time, using few seconds of reference audio.
- stars
- 7.7K
- forks
- 755
- watching
- 89
Similar projects
Found by comparing what the projects do, not just their names.
Voice Generator
An AI system for generating human-like voices from text inputs, using deep learning techniques and pre-trained models.
neonbjb/tortoise-tts13.4K
TTS system
An open-source text-to-speech system trained with high-quality audio capabilities
Speech synthesizer
A deep learning model for generating human-like speech
coqui-ai/tts36.1K
Speech generator
A deep learning toolkit for generating human-like speech from text
mozilla/tts9.5K
Text-to-Speech Library
An open-source project providing a suite of deep learning models and tools for advanced text-to-speech synthesis.
suno-ai/bark36.4K
Audio generator
A text-to-audio model that generates realistic speech and other audio
Speech Synthesizer
Real-time speech synthesis using state-of-the-art architectures
Prompt generator
A tool for automating the process of generating and ranking effective prompts for AI models like GPT-4, GPT-3.5-Turbo, or Claude 3 Opus.
Speech synthesizer
A deep learning-based speech synthesis model that generates natural-sounding audio with controlled prosody.
Audio generator
A deep learning library for generating high-quality audio
Text Generation Toolkit
A toolkit for deploying and serving Large Language Models (LLMs) for high-performance text generation
Code assistant
A plugin for Neovim that integrates with the ChatGPT API to generate natural language responses and assist with coding tasks.
TTS system
Develops an end-to-end text-to-speech system that generates more natural audio than existing models
Conversational framework
A modular framework for building conversational AI applications with real-time voice and multimodal interactions.
Audio Generator
A Python-based audio generation tool that can produce speech, sound effects, music, and more, using text as input or guided by user description.