opencompass

LLM evaluator

An LLM evaluation platform supporting various models and datasets

OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.

GitHub

4k stars
26 watching
457 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list

benchmarkchatgptevaluationlarge-language-modelllama2llama3llmopenai

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
imoneoi/openchatFine-tuned language models trained on mixed-quality data5,273
eleutherai/lm-evaluation-harnessProvides a unified framework to test generative language models on various evaluation tasks.7,200
open-compass/vlmevalkitAn evaluation toolkit for large vision-language models1,514
ianarawjo/chainforgeAn environment for battle-testing prompts to Large Language Models (LLMs) to evaluate response quality and performance.2,413
alpha-vllm/llama2-accessoryAn open-source toolkit for pretraining and fine-tuning large language models2,732
traceloop/openllmetryA set of extensions built on top of OpenTelemetry to provide observability for large language model applications.5,188
fittentech/openllama-chineseA Chinese language large language model built from OpenLLaMA and fine-tuned on various datasets for multilingual text generation.65
openlmlab/openchinesellamaAn incremental pre-trained Chinese large language model based on the LLaMA-7B model234
brexhq/prompt-engineeringGuides software developers on how to effectively use and build systems around Large Language Models like GPT-4.8,487
langfuse/langfuseAn integrated development platform for large language models (LLMs) that provides observability, analytics, and management tools.7,123
openbmb/toolbenchA platform for training, serving, and evaluating large language models to enable tool use capability4,888
opengvlab/llama-adapterAn implementation of a method for fine-tuning language models to follow instructions with high efficiency and accuracy5,775
open-compass/lawbenchEvaluates the legal knowledge of large language models using a custom benchmarking framework.273
thunlp/openpromptA flexible framework for adapting pre-trained language models to downstream NLP tasks using textual templates4,398
openai-translator/openai-translatorA multi-platform translator and text processing tool leveraging ChatGPT API24,004