Pengi

Audio Model

An Audio Language Model framework that uses transfer learning to generate text from audio inputs

An Audio Language model for Audio Tasks

GitHub

295 stars
14 watching
17 forks
Language: Python
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
yuangongnd/ltuAn audio and speech large language model implementation with pre-trained models, datasets, and inference options396
balavenkatesh3322/audio-pretrained-modelA collection of pre-trained audio and speech models for various applications183
qwenlm/qwen-audioA multimodal audio language model developed by Alibaba Cloud that supports various tasks and languages1,515
ibm/max-audio-classifierIdentifies sounds in short audio clips using machine learning and PCA transformation154
yongxuustc/dcase2017_task4_cvsspA system for audio classification and detection using machine learning models4
elanmart/psmmAn implementation of a neural network model for character-level language modeling.50
qwenlm/qwen2-audioAn audio-language model that can analyze or respond to speech instructions based on audio input1,306
awni/speechA PyTorch implementation of end-to-end speech recognition models.756
openai/finetune-transformer-lmThis project provides code and model for improving language understanding through generative pre-training using a transformer-based architecture.2,167
jordipons/music-audio-tagging-at-scale-modelsResearch on end-to-end learning for music audio tagging using large datasets and different front-end paradigms.149
microsoft/mpnetDevelops a method for pre-training language understanding models by combining masked and permuted techniques, and provides code for implementation and fine-tuning.288
jthorborg/apeAn Audio Programming Environment with support for AU and DSP plugins14
keunwoochoi/auralisationReconstructs audio features learned by convolutional neural networks into audible sounds42
cpjku/madmomA Python audio signal processing library used in music information retrieval tasks.1,366
ynop/audiomateA Python library for handling audio datasets, providing tools for accessing, manipulating, and preparing data for machine learning tasks.133