Video-MME

Video analysis benchmark

Comprehensive benchmark for evaluating multi-modal large language models on video analysis tasks

✨✨Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

GitHub

422 stars
5 watching
17 forks
last commit: almost 2 years ago
large-language-modelslarge-vision-language-modelsmmemultimodal-large-language-modelsvideovideo-mme

Related projects:

RepositoryDescriptionStars
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84
pku-yuangroup/video-benchEvaluates and benchmarks large language models' video understanding capabilities121
rese1f/moviechatDevelops a method for long video understanding by optimizing memory usage550
tsb0601/mmvpAn evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.296
felixgithub2017/mmcuMeasures the understanding of massive multitask Chinese datasets using large language models87
boheumd/ma-lmmThis project develops an AI model for long-term video understanding254
yfzhang114/mme-realworldA multimodal large language model benchmark designed to simulate real-world challenges and measure the performance of such models in practical scenarios.86
pku-yuangroup/chronomagic-benchProvides a benchmarking framework for evaluating the quality of text-to-video generation models191
huaizhengzhang/awsome-deep-learning-for-video-analysisA collection of resources and tools for video analysis using deep learning and multi-modal learning techniques.767
cmmmu-benchmark/cmmmuA benchmark for evaluating the performance of multimodal question answering models on diverse domains and data types46
mltframework/mltA multimedia framework designed for video editing, providing tools and libraries for audio and video processing.1,522
damo-nlp-sg/m3examA benchmark for evaluating large language models in multiple languages and formats93
gabeur/mmtDevelops a cross-modal architecture for video retrieval by combining multiple types of features from videos and text259
chenllliang/mmevalproA benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.22
laomao0/binSoftware to interpolate blurry video frames and enhance image quality209