FlagEval

Model evaluation framework

An evaluation toolkit and platform for assessing large models in various domains

FlagEval is an evaluation toolkit for AI large foundation models.

GitHub

307 stars
13 watching
27 forks
Language: Python
last commit: about 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
flagai-open/aquila2Provides pre-trained language models and tools for fine-tuning and evaluation439
allenai/olmo-evalA framework for evaluating language models on NLP tasks326
huggingface/lightevalAn all-in-one toolkit for evaluating Large Language Models (LLMs) across multiple backends.879
open-evals/evalsA framework for evaluating OpenAI models and an open-source registry of benchmarks.19
modelscope/evalscopeA framework for efficiently evaluating and benchmarking large models308
huggingface/evaluateAn evaluation framework for machine learning models and datasets, providing standardized metrics and tools for comparing model performance.2,063
openai/simple-evalsEvaluates language models using standardized benchmarks and prompting techniques.2,059
psycoy/mixevalAn evaluation suite and dynamic data release platform for large language models230
stanford-crfm/helmA framework to evaluate and compare language models by analyzing their performance on various tasks1,981
baaivision/emuA multimodal generative model framework1,672
chenllliang/mmevalproA benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.22
declare-lab/instruct-evalAn evaluation framework for large language models trained with instruction tuning methods535
aiverify-foundation/llm-evals-catalogueA collaborative catalogue of LLM evaluation frameworks and papers13
ukgovernmentbeis/inspect_aiA framework for evaluating large language models669
maluuba/nlg-evalA toolset for evaluating and comparing natural language generation models1,350