transfusion-asr

Speech Transcription Tool

An ASR project that uses diffusion models to transcribe speech

Transcribing Speech with Multinomial Diffusion, training code and models.

GitHub

76 stars
8 watching
5 forks
Language: Python
last commit: almost 3 years ago
Linked from 1 awesome list

asrbinomial-distributiondiffusiondiscrete-diffusionpytorchspeech-recognition

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
shashikg/whispers2tAn optimized speech-to-text pipeline designed to improve inference speed and accuracy330
bnosac/audio.whisperProvides an R interface to the Whisper Automatic Speech Recognition model119
linto-ai/whisper-timestampedAn extension to the Whisper speech recognition model that adds word-level timestamps and confidence scores.2,121
collabora/whisperliveAn implementation of Whisper's speech-to-text functionality in a real-time transcription application2,186
arthurfdlr/whisper-youtubeTranscribes Youtube videos using OpenAI's Whisper speech recognition model369
matlab-deep-learning/deepspeechEnables speech-to-text transcription using a pre-trained Deep Speech model in MATLAB.7
ytsvetko/str2ipaA tool for phonetic transcription of languages with close-to-phonetic writing systems10
neso613/asr_tfliteProvides pre-trained ASR models for efficient inference using TFLite11
dodohow1011/speechadvreprogramDeveloping low-resource speech command recognition systems using adversarial reprogramming and transfer learning18
birch-san/diffusersA toolkit for creating and manipulating state-of-the-art diffusion models in PyTorch8
langtech/transcriberAn online transcription tool for a specific application, allowing users to input audio or video and receive a written text summary2
srijith-rkr/kaust-whisper-adapterA tool for fine-tuning the OpenAI Whisper speech recognition model using residual adapters and parameter-efficient learning methods.32
awni/speechA PyTorch implementation of end-to-end speech recognition models.756
r3gm/sonitranslateSoftware that allows video translation with synchronized audio, utilizing speech-to-text and text-to-speech technologies.924
mybigday/whisper.rnA React Native binding of Whisper's automatic speech recognition model408