CogVLM
by THUDM
a state-of-the-art-level open visual language model | 多模态预训练模型
AI summary
Visual Language Model
Develops a state-of-the-art visual language model with applications in image understanding and dialogue systems.
- stars
- 6.2K
- forks
- 419
- watching
- 68
Similar projects
Found by comparing what the projects do, not just their names.
thudm/cogvideo9.8K
Video generator
Generates videos from text and images using large language models
thudm/glm3.2K
Language Model
A general-purpose language model pre-trained with an autoregressive blank-filling objective and designed for various natural language understanding and generation tasks.
Instruction-following model tuner
An implementation of a method for fine-tuning language models to follow instructions with high efficiency and accuracy
Multimodal model builder
Develops large language models capable of processing multiple data types and modalities
qwenlm/qwen-vl5.2K
Large vision language model
A large vision language model with improved image reasoning and text recognition capabilities, suitable for various multimodal tasks
openbmb/minicpm-v12.9K
Multimodal LLM
A multimodal language model designed to understand images, videos, and text inputs and generate high-quality text outputs.
Evaluation framework
An evaluation toolkit for large vision-language models
haotian-liu/llava20.7K
Visual Instruction System
A system that uses large language and vision models to generate and process visual instructions
replicate/cog8.2K
Model deployer
A tool for packaging and deploying machine learning models in a standard, production-ready container environment.
thudm/glm-130b7.7K
Bilingual Language Model
An open-source implementation of a large bilingual language model pre-trained on vast amounts of text data.
Visual decoder
A large language model designed to process and generate visual information
LLM toolkit
An open-source toolkit for pretraining and fine-tuning large language models
Tool learning platform
A platform for training, serving, and evaluating large language models to enable tool use capability
antvis/g212.2K
Visualization framework
A visualization grammar that enables rapid creation of data-driven visualizations with concise declarations and infers complex details.
open-mmlab/mmcv5.9K
Computer Vision Library
Provides a foundational library for computer vision research and training deep learning models with high-quality implementation of common CPU and CUDA ops.