ALLaVA
Harnessing 1.4M GPT4V-synthesized Data for A Lite Vision-Language Model
AI summary
Vision-Language Model Dataset
A collection of datasets and models designed to support the training of lite vision-language models.
- stars
- 249
- forks
- 9
- watching
- 11
Similar projects
Found by comparing what the projects do, not just their names.
Image processor
A system for scaling large language models to process and understand visual information from multiple images efficiently.
Vision-Language Model
A PyTorch implementation of an encoder-free vision-language model that can be fine-tuned for various tasks and modalities
nvlabs/prismer1.3K
Vision-Language Model
A deep learning framework for training multi-modal models with vision and language capabilities.
Long context transfer
An open-source project that enables the transfer of language understanding to vision capabilities through long context processing.
Vision-Language Model
A multimodal AI model that enables real-world vision-language understanding applications
Model evaluator
Evaluates and compares the performance of multimodal large language models on various tasks
Visual Prompt Model
A system designed to enable large multimodal models to understand arbitrary visual prompts
Image segmentation tool
A system that uses large language models to generate segmentation masks for images based on complex queries and world knowledge.
Image dataset generator
Scripts to generate datasets for an image generation task using Generative Adversarial Networks and deep learning techniques
LLM
An open-source implementation of a vision-language instructed large language model
Vision language model trainer
An annotated preference dataset and training framework for improving large vision language models.
Vision-Language Model Framework
Implementing a unified modal learning framework for generative vision-language models
Medical LLM
Developing a large language model for medical consultations by combining distilled and real-world data to improve doctor-patient interactions
Model trainer
A platform for training and deploying large language and vision models that can use tools to perform tasks
jy0205/lavit544
Visual understanding and generation framework
A unified framework for training large language models to understand and generate visual content