SALMONN

Audio perceptron

A large language model enabling speech, audio event perception and music inputs to achieve multilingual capabilities

SALMONN: Speech Audio Language Music Open Neural Network

GitHub

1k stars
28 watching
85 forks
Language: Python
last commit: almost 2 years ago
audioaudio-processingbytedanceiclr2024icml-2024large-language-modelsmulti-modalmusicresearchspeechspeech-recognitiontsinghua-university

Related projects:

RepositoryDescriptionStars
keunwoochoi/auralisationReconstructs audio features learned by convolutional neural networks into audible sounds42
soerenab/audiomnistThis project provides an implementation of a deep learning framework to classify audio signals and offers insights into the model's decision-making process using Explainable Artificial Intelligence (AI) techniques.351
soroushmehr/samplernn_iclr2017An unconditional end-to-end neural audio generation model utilizing a recurrent neural network architecture.537
ibm/max-audio-classifierIdentifies sounds in short audio clips using machine learning and PCA transformation154
kinwaicheuk/nnaudioAn audio processing toolkit using PyTorch convolutional neural networks to generate spectrograms from raw audio data1,036
yuangongnd/ltuAn audio and speech large language model implementation with pre-trained models, datasets, and inference options396
yongxuustc/dcase2017_task4_cvsspA system for audio classification and detection using machine learning models4
balavenkatesh3322/audio-pretrained-modelA collection of pre-trained audio and speech models for various applications183
drscotthawley/audio-classifier-keras-cnnAn audio classification system using a convolutional neural network to classify audio data160
ksw0306/clarinetAn implementation of a neural network-based vocoder using parallel-wavenet architecture and autoregressive flow290
deepsound-project/samplernn-pytorchAn implementation of an audio generation model using PyTorch290
dodohow1011/speechadvreprogramDeveloping low-resource speech command recognition systems using adversarial reprogramming and transfer learning18
xidongwu/d-auprcProvides an implementation of a specific algorithm used in audio signal processing0
mlachmish/musicgenreclassificationClassify music genre from a 10-second sound stream using a neural network.565
microsoft/pengiAn Audio Language Model framework that uses transfer learning to generate text from audio inputs295