LanguageBind
【ICLR 2024🔥】 Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
AI summary
Multimodal alignment model
Extending pretraining models to handle multiple modalities by aligning language and video representations
- stars
- 751
- forks
- 52
- watching
- 15
Similar projects
Found by comparing what the projects do, not just their names.
Model aligner
Aligns large multimodal models with human intentions and values using various algorithms and fine-tuning methods.
Video benchmarking toolkit
Evaluates and benchmarks large language models' video understanding capabilities
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
Vision-Language Learning Model
Develops and trains models for vision-language learning with decoupled language pre-training
Region-of-Interest Training
Training and deploying large language models on computer vision tasks using region-of-interest inputs
Multimodal LLM
A multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation
Language Model
Large-scale language model with improved performance on NLP tasks through distributed training and efficient data processing
Attention calibrator
This project proposes a novel method for calibrating attention distributions in multimodal models to improve contextualized representations of image-text pairs.
Model alignment
This project demonstrates the effectiveness of reinforcement learning from human feedback (RLHF) in improving small language models like GPT-2.
Mixture of Experts Model
A large vision-language model using a mixture-of-experts architecture to improve performance on multi-modal learning tasks
Visual unification framework
A framework for unified visual representation in image and video understanding models, enabling efficient training of large language models on multimodal data.
Multimodal LLM
A multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks
Conversational AI framework
Enables larger language models to generate multi-turn multimodal instruction-response conversations from image-caption pairs with minimal annotations.
Chinese language model
Trains a large Chinese language model on massive data and provides a pre-trained model for downstream tasks
Multimodal learner
Develops a multimodal vision-language model to enable machines to understand complex relationships between instructions and images in various tasks.