AMBER
An LLM-free Multi-dimensional Benchmark for Multi-modal Hallucination Evaluation
AI summary
MLLM benchmark
An LLM-free benchmark suite for evaluating MLLMs' hallucination capabilities in various tasks and dimensions
- stars
- 98
- forks
- 2
- watching
- 1
Similar projects
Found by comparing what the projects do, not just their names.
Hallucination tester
Evaluates and mitigates hallucinations in multimodal large language models
Hallucination detector
A framework for detecting hallucinations in large language models
LM Benchmark
A benchmark for evaluating large language models in multiple languages and formats
Model Comparison
Compares performance of large language models on generating coherent summaries from short documents
Model evaluator
Evaluates and compares the performance of multimodal large language models on various tasks
Visual Model Benchmark
An open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models
Multimodal benchmarking
Evaluates and benchmarks multimodal language models' ability to process visual, acoustic, and textual inputs simultaneously.
Benchmark
An image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy
Multimodal LLM test suite
A benchmark for evaluating large language models' ability to process multimodal input
ML Model Benchmarker
Provides a benchmarking framework and dataset for evaluating the performance of large language models in text-to-image tasks
LLM exploration
Exploring various LLMs and their applications in natural language processing and related areas
Hallucination corrector
A method to correct hallucinations in multimodal large language models without requiring retraining
Machine Learning Logger
A lightweight logger for machine learning experiments
Multimodal LLM
A multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks
ML benchmark
A benchmark for evaluating machine learning algorithms' performance on large datasets