ltu

Audio Model

An audio and speech large language model implementation with pre-trained models, datasets, and inference options

Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".

GitHub

396 stars
15 watching
38 forks
Language: Python
last commit: over 2 years ago
audioaudio-processingdeep-learninglarge-language-modelsspeech-recognition

Related projects:

RepositoryDescriptionStars
microsoft/pengiAn Audio Language Model framework that uses transfer learning to generate text from audio inputs295
balavenkatesh3322/audio-pretrained-modelA collection of pre-trained audio and speech models for various applications183
shawn-ieitsystems/yuan-1.0Large-scale language model with improved performance on NLP tasks through distributed training and efficient data processing591
brightmart/xlnet_zhTrains a large Chinese language model on massive data and provides a pre-trained model for downstream tasks230
ymcui/lertA pre-trained language model designed to leverage linguistic features and outperform comparable baselines on Chinese natural language understanding tasks.202
ieit-yuan/yuan2.0-m32A high-performance language model designed to excel in tasks like natural language understanding, mathematical computation, and code generation182
yunwentechnology/unilmThis project provides pre-trained models and tools for natural language understanding (NLU) and generation (NLG) tasks in Chinese.439
qwenlm/qwen-audioA multimodal audio language model developed by Alibaba Cloud that supports various tasks and languages1,515
yuangongnd/whisper-atAn audio processing model that adds audio event tagging capabilities to an existing speech recognition system with minimal additional computational cost.343
bytedance/salmonnA large language model enabling speech, audio event perception and music inputs to achieve multilingual capabilities1,091
tencent/tencent-hunyuan-largeThis project makes a large language model accessible for research and development1,245
thu-coai/opdA large-scale pre-trained dialogue model for Chinese language74
renshuhuai-andy/timechatA large language model designed to understand long videos by binding visual content with timestamps and producing video token sequences of varying lengths.314
baai-wudao/modelA repository of pre-trained language models for various tasks and domains.121
qwenlm/qwen2-audioAn audio-language model that can analyze or respond to speech instructions based on audio input1,306