langtest

Model Tester

A tool for testing and evaluating large language models with a focus on AI safety and model assessment.

Deliver safe & effective language models

GitHub

506 stars
10 watching
41 forks
Language: Python
last commit: almost 2 years ago
ai-safetyai-testingartificial-intelligencebenchmark-frameworkbenchmarksethics-in-ailarge-language-modelsllmllm-as-evaluatorllm-evaluation-toolkitllm-testllm-testingml-safetyml-testingmlopsmodel-assessmentnlpresponsible-aitrustworthy-ai

Related projects:

RepositoryDescriptionStars
howiehwong/trustllmA toolkit for assessing trustworthiness in large language models491
aiplanethub/beyondllmAn open-source toolkit for building and evaluating large language models267
declare-lab/instruct-evalAn evaluation framework for large language models trained with instruction tuning methods535
neulab/explainaboardAn interactive tool to analyze and compare the performance of natural language processing models362
vhellendoorn/code-lmsA guide to using pre-trained large language models in source code analysis and generation1,789
comet-ml/opikA platform for evaluating and testing large language models (LLMs) during development and production.2,588
freedomintelligence/mllm-benchEvaluates and compares the performance of multimodal large language models on various tasks56
innogames/ltcA tool for managing load tests and analyzing performance results200
qcri/llmebenchA benchmarking framework for large language models81
maluuba/nlg-evalA toolset for evaluating and comparing natural language generation models1,350
openlmlab/gaokao-benchAn evaluation framework using Chinese high school examination questions to assess large language model capabilities565
flagai-open/aquila2Provides pre-trained language models and tools for fine-tuning and evaluation439
bilibili/index-1.9bA lightweight, multilingual language model with a long context length920
01-ai/yiA series of large language models trained from scratch to excel in multiple NLP tasks7,743
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322