IDA-VLM
by jiyt17
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
AI summary
Identity-aware video model
An open-source project that aims to improve large vision-language models by integrating identity-aware capabilities and utilizing visual instruction tuning data for movie understanding
- stars
- 26
- forks
- 0
- watching
- 4
Similar projects
Found by comparing what the projects do, not just their names.
Image understanding model
A large multi-modal model developed using the Llama3 language model, designed to improve image understanding capabilities.
Video understanding model
This project develops an AI model for long-term video understanding
Vision Language Model
An implementation of a vision language model designed for mobile devices, utilizing a lightweight downsample projector and pre-trained language models.
Language model toolkit
A guide to using pre-trained large language models in source code analysis and generation
Visual decoder
A large language model designed to process and generate visual information
Vision-Language Model
A multimodal AI model that enables real-world vision-language understanding applications
Video image processor
An image-based language model that uses large language models to generate visual and text features from videos
aidc-ai/ovis575
Multimodal aligner
An MLLM architecture designed to align visual and textual embeddings through structural alignment
Model interpreter
A toolset for understanding and interpreting complex machine learning models
Image descriptor model
Improves large vision-language models' ability to accurately describe images by combining global and local attention mechanisms.
Language Model
A high-performance language model designed to excel in tasks like natural language understanding, mathematical computation, and code generation
Long context transfer
An open-source project that enables the transfer of language understanding to vision capabilities through long context processing.
Driving Model
An autonomous driving project exploring the capabilities of a visual-language model in understanding complex driving scenes and making decisions
Visual Model Benchmark
An open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models
Language Model
An implementation of a neural network model for character-level language modeling.