SliME
by yfzhang114
✨✨Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models
AI summary
Multimodal model developer
Develops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.
- stars
- 143
- forks
- 7
- watching
- 4
Similar projects
Found by comparing what the projects do, not just their names.
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Model debiasing
Debiasing techniques to minimize hallucinations in large visual language models
Multimodal model
A large multimodal model for visual question answering, trained on a dataset of 2.1B image-text pairs and 8.2M instruction sequences.
Real-world challenge simulator
A multimodal large language model benchmark designed to simulate real-world challenges and measure the performance of such models in practical scenarios.
Multilingual Model
Develops and publishes large multilingual language models with advanced mixing-of-experts architecture.
openbmb/viscpm1.1K
Multimodal Models
A family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages
Multimodal LLM
A multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks
LLM
A multilingual large language model developed by XVERSE Technology Inc.
Multilingual models
Large language models designed to perform well in multiple languages and address performance issues with current multilingual models.
Multimodal LLM
An open-source multilingual large language model designed to understand and generate content across diverse languages and cultural contexts
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
lxtgh/omg-seg1.3K
Visual Model
Develops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.
3D LLM
Developing a Large Language Model capable of processing 3D representations as inputs
Chinese understanding benchmark
Measures the understanding of massive multitask Chinese datasets using large language models