InternVL

Multimodal model builder

Develops large language models capable of processing multiple data types and modalities

[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型

GitHub

6k stars
57 watching
493 forks
Language: Python
last commit: almost 2 years ago
gptgpt-4ogpt-4vimage-classificationimage-text-retrievalllmmulti-modalsemantic-segmentationvideo-classificationvision-language-modelvit-22bvit-6b

Related projects:

RepositoryDescriptionStars
openbmb/minicpm-vA multimodal language model designed to understand images, videos, and text inputs and generate high-quality text outputs.12,870
internlm/lmdeployA toolkit for optimizing and serving large language models4,854
thudm/cogvlmDevelops a state-of-the-art visual language model with applications in image understanding and dialogue systems.6,182
internlm/internlmA collection of large language models designed to improve reasoning and tool use capabilities in chatbots.6,572
opengvlab/internvideoDevelops general video foundation models and related datasets for multimodal understanding and generation through generative and discriminative learning.1,467
vision-cair/minigpt-4Enabling vision-language understanding by fine-tuning large language models on visual data.25,490
memochou1993/gpt-ai-assistantAn AI-powered chat application leveraging OpenAI models and LINE APIs for conversational interfaces.7,491
open-mmlab/mmaction2A comprehensive video understanding toolbox and benchmark with modular design, supporting various tasks such as action recognition, localization, and retrieval.4,360
openai/gpt-2A repository providing code and models for research into language modeling and multitask learning22,644
open-mmlab/mmcvProvides a foundational library for computer vision research and training deep learning models with high-quality implementation of common CPU and CUDA ops.5,948
internlm/internlm-xcomposerA comprehensive multimodal system for long-term streaming video and audio interactions with capabilities including text-image comprehension and composition2,616
open-compass/opencompassAn LLM evaluation platform supporting various models and datasets4,295
opengvlab/llama-adapterAn implementation of a method for fine-tuning language models to follow instructions with high efficiency and accuracy5,775
doubiiu/dynamicrafterThis project generates animated videos from open-domain images by leveraging pre-trained video diffusion priors.2,668
thudm/glm-4A large language model designed for multilingual and multimodal chat applications with advanced features such as long-text reasoning and high-performance inference.5,525