inspect_ai

Model inspector

A framework for evaluating large language models

Inspect: A framework for large language model evaluations

GitHub

669 stars
9 watching
135 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
declare-lab/instruct-evalAn evaluation framework for large language models trained with instruction tuning methods535
allenai/olmo-evalA framework for evaluating language models on NLP tasks326
openai/simple-evalsEvaluates language models using standardized benchmarks and prompting techniques.2,059
klen/pylamaAutomates code quality checks for Python programs1,049
flageval-baai/flagevalAn evaluation toolkit and platform for assessing large models in various domains307
ruixiangcui/agievalEvaluates foundation models on human-centric tasks with diverse exams and question types714
ilevkivskyi/typing_inspectProvides utilities for inspecting and analyzing Python types at runtime352
johnsnowlabs/langtestA tool for testing and evaluating large language models with a focus on AI safety and model assessment.506
modelscope/evalscopeA framework for efficiently evaluating and benchmarking large models308
open-compass/lawbenchEvaluates the legal knowledge of large language models using a custom benchmarking framework.273
openlmlab/gaokao-benchAn evaluation framework using Chinese high school examination questions to assess large language model capabilities565
huggingface/evaluateAn evaluation framework for machine learning models and datasets, providing standardized metrics and tools for comparing model performance.2,063
h2oai/mli-resourcesProvides tools and techniques for interpreting machine learning models483
flagai-open/aquila2Provides pre-trained language models and tools for fine-tuning and evaluation439
xverse-ai/xverse-moe-a36bDevelops and publishes large multilingual language models with advanced mixing-of-experts architecture.37