FFTNet

Text-to-Speech Model

An implementation of a deep learning-based text-to-speech model using the FFTNet architecture.

FFTNet vocoder implementation

GitHub

81 stars
15 watching
8 forks
Language: Jupyter Notebook
last commit: almost 8 years ago
deep-learningfftnetpytorchtext2speechvocoder

Related projects:

RepositoryDescriptionStars
lifeiteng/vall-eA PyTorch implementation of a text-to-speech synthesizer based on large language models2,062
soobinseo/tacotron-pytorchA PyTorch implementation of an end-to-end text-to-speech synthesis model.207
kefirski/bytenetA Pytorch implementation of a neural network model for machine translation47
awni/speechA PyTorch implementation of end-to-end speech recognition models.756
r9y9/deepvoice3_pytorchAn implementation of text-to-speech synthesis using convolutional neural networks in PyTorch1,970
openai/finetune-transformer-lmThis project provides code and model for improving language understanding through generative pre-training using a transformer-based architecture.2,167
isht7/pytorch-deeplab-resnetA deep learning model implementation of the DeepLab ResNet architecture for image segmentation tasks.602
l0sg/relational-rnn-pytorchAn implementation of DeepMind's Relational Recurrent Neural Networks (Santoro et al. 2018) in PyTorch for word language modeling245
gram-ai/radio-transformer-networksAn implementation of a machine learning-based communications system using deep learning techniques.127
eromera/erfnetA toolbox for training and evaluating real-time semantic segmentation networks using Torch library.120
matlab-deep-learning/deepspeechEnables speech-to-text transcription using a pre-trained Deep Speech model in MATLAB.7
randl/shufflenetv2-pytorchAn implementation of a lightweight convolutional neural network architecture for mobile devices191
yuangongnd/ltuAn audio and speech large language model implementation with pre-trained models, datasets, and inference options396
erogol/seg-torchCustom image segmentation implementation using deep learning with Lua and Torch37
taoxugit/attnganReproduces text-to-image generation with attentional generative adversarial networks.1,343