TempCompass

Video understanding tester

A tool to evaluate video language models' ability to understand and describe video content

[ACL 2024 Findings] "TempCompass: Do Video LLMs Really Understand Videos?", Yuanxin Liu, Shicheng Li, Yi Liu, Yuxiang Wang, Shuhuai Ren, Lei Li, Sishuo Chen, Xu Sun, Lu Hou

GitHub

91 stars
4 watching
2 forks
Language: Python
last commit: almost 2 years ago
evaluationtemporal-perceptionvideo-llms

Related projects:

RepositoryDescriptionStars
mit-han-lab/temporal-shift-moduleDevelops a video analysis module with efficient temporal processing capabilities2,078
lxtgh/omg-segDevelops an end-to-end model for multiple visual perception and reasoning tasks using a single encoder, decoder, and large language model.1,336
pku-yuangroup/video-benchEvaluates and benchmarks large language models' video understanding capabilities121
boheumd/ma-lmmThis project develops an AI model for long-term video understanding254
dcdmllm/momentorA video Large Language Model designed for fine-grained comprehension and localization in videos with a custom Temporal Perception Module for improved temporal modeling58
renshuhuai-andy/timechatA large language model designed to understand long videos by binding visual content with timestamps and producing video token sequences of varying lengths.314
antoine77340/howto100mProvides code and tools for learning joint text-video embeddings using the HowTo100M dataset254
dvlab-research/llama-vidAn image-based language model that uses large language models to generate visual and text features from videos748
poyro/poyroAn extension of Vitest for testing LLM applications using local language models31
researchmm/sttnProposes a deep learning model to fill missing regions in video frames and generate completed videos480
tsb0601/mmvpAn evaluation framework for multimodal language models' visual capabilities using image and question benchmarks.296
johnsnowlabs/langtestA tool for testing and evaluating large language models with a focus on AI safety and model assessment.506
krrishdholakia/betterpromptAn API for evaluating the quality of text prompts used in Large Language Models (LLMs) based on perplexity estimation43
huangb23/vtimellmA PyTorch-based Video LLM designed to understand and reason about video moments in terms of time boundaries.231
ray-project/llmperfA tool for evaluating the performance of large language model APIs678