Monkey
by Yuliang-Liu
【CVPR 2024 Highlight】Monkey (LMM): Image Resolution and Text Label Are Important Things for Large Multi-modal Models
AI summary
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
- stars
- 1.8K
- forks
- 132
- watching
- 22
Similar projects
Found by comparing what the projects do, not just their names.
Multimodal model framework
A framework for grounding language models to images and handling multimodal inputs and outputs
Multimodal LLM
A multi-modal language model that integrates image, video, audio, and text data to improve language understanding and generation
OCR Benchmark
An evaluation benchmark for OCR capabilities in large multmodal models.
Multimodal LLM
A multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks
openbmb/viscpm1.1K
Multimodal Models
A family of large multimodal models supporting multimodal conversational capabilities and text-to-image generation in multiple languages
Multimodal model developer
Develops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
Multimodal LLM framework
A framework for training GPT4-style language models with multimodal inputs using large datasets and pre-trained models
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics
Multimodal Model Builder
A framework to build versatile Multimodal Large Language Models with synergistic comprehension and creation capabilities
Multimodal LLM Framework
A framework that enables large language models to process and understand multimodal inputs from various sources such as images and speech.
Multimodal LLM
An implementation of a multimodal language model with capabilities for comprehension and generation
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
Multimodal conversational model
An end-to-end trained model capable of generating natural language responses integrated with object segmentation masks for interactive visual conversations
yuxie11/r2d2157
Vision-Language Framework
A framework for large-scale cross-modal benchmarks and vision-language tasks in Chinese