AGIEval
by ruixiangcui
AI summary
Exam evaluator
Evaluates foundation models on human-centric tasks with diverse exams and question types
- stars
- 714
- forks
- 48
- watching
- 9
Similar projects
Found by comparing what the projects do, not just their names.
Model Evaluator
Evaluates language models using standardized benchmarks and prompting techniques.
Chart eval
An evaluation suite for assessing chart understanding in multimodal large language models.
Evaluation pipeline
A framework for evaluating language models on NLP tasks
Prompt evaluator
An API for evaluating the quality of text prompts used in Large Language Models (LLMs) based on perplexity estimation
hkust-nlp/ceval1.7K
Evaluation suite
An evaluation suite providing multiple-choice questions for foundation models in various disciplines, with tools for assessing model performance.
Segmentation evaluator
Evaluates segmentation performance in medical imaging using multiple metrics
cloud-cv/evalai1.8K
Benchmarking tool
A platform for comparing and evaluating AI and machine learning algorithms at scale
Evaluator
An automatic evaluation tool for large language models
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
QA evaluation tool
Automated evaluation of language models for question answering tasks
Skill evaluator
An online platform offering innovative evaluation and certification of digital skills
Multimodal model evaluator
Evaluating and improving large multimodal models through in-context learning
Word vector evaluator
A set of Python scripts for evaluating word vectors on various tasks and comparing similarity between words.
OCR evaluator
A collection of tools and utilities for evaluating the performance and quality of OCR output
maja42/goval160
Expression evaluator
A Go library for evaluating arbitrary arithmetic, string, and logic expressions with support for variables and custom functions.