IDA-VLM

Identity-aware video model

An open-source project that aims to improve large vision-language models by integrating identity-aware capabilities and utilizing visual instruction tuning data for movie understanding

IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model

GitHub

26 stars
4 watching
0 forks
Language: Python
last commit: almost 2 years ago

Related projects:

RepositoryDescriptionStars
360cvgroup/360vlA large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.32
boheumd/ma-lmmThis project develops an AI model for long-term video understanding254
meituan-automl/mobilevlmAn implementation of a vision language model designed for mobile devices, utilizing a lightweight downsample projector and pre-trained language models.1,076
vhellendoorn/code-lmsA guide to using pre-trained large language models in source code analysis and generation1,789
opengvlab/visionllmA large language model designed to process and generate visual information956
deepseek-ai/deepseek-vlA multimodal AI model that enables real-world vision-language understanding applications2,145
dvlab-research/llama-vidAn image-based language model that uses large language models to generate visual and text features from videos748
aidc-ai/ovisAn MLLM architecture designed to align visual and textual embeddings through structural alignment575
mayer79/flashlightA toolset for understanding and interpreting complex machine learning models22
lackel/aglaImproves large vision-language models' ability to accurately describe images by combining global and local attention mechanisms.18
ieit-yuan/yuan2.0-m32A high-performance language model designed to excel in tasks like natural language understanding, mathematical computation, and code generation182
evolvinglmms-lab/longvaAn open-source project that enables the transfer of language understanding to vision capabilities through long context processing.347
pjlab-adg/gpt4v-ad-explorationAn autonomous driving project exploring the capabilities of a visual-language model in understanding complex driving scenes and making decisions288
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84
elanmart/psmmAn implementation of a neural network model for character-level language modeling.50