jury

NLP evaluator

A comprehensive toolkit for evaluating NLP experiments offering automated metrics and efficient computation.

Comprehensive NLP Evaluation System

GitHub

187 stars
5 watching
20 forks
Language: Python
last commit: about 2 years ago
Linked from 1 awesome list

datasetsevaluateevaluationhuggingfacemachine-learningmetricsnatural-language-processingnlpnlp-evaluationpythonpytorchtransformers

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
maluuba/nlg-evalA toolset for evaluating and comparing natural language generation models1,350
huggingface/evaluateAn evaluation framework for machine learning models and datasets, providing standardized metrics and tools for comparing model performance.2,063
allenai/olmo-evalA framework for evaluating language models on NLP tasks326
nullne/evaluatorAn expression evaluator library written in Go.41
openai/simple-evalsEvaluates language models using standardized benchmarks and prompting techniques.2,059
olical/conjureAn interactive environment for evaluating code within a running program.1,806
open-compass/lawbenchEvaluates the legal knowledge of large language models using a custom benchmarking framework.273
tatsu-lab/alpaca_evalAn automatic evaluation tool for large language models1,568
princeton-nlp/charxivAn evaluation suite for assessing chart understanding in multimodal large language models.85
ermlab/polish-word-embeddings-reviewAn evaluation framework for Polish word embeddings prepared by various research groups using analogy tasks.4
huggingface/lightevalAn all-in-one toolkit for evaluating Large Language Models (LLMs) across multiple backends.879
lartpang/pysodevaltoolkitA comprehensive Python toolbox for evaluating salient object detection and camouflaged object detection tasks168
eddieantonio/ocrevalA collection of tools and utilities for evaluating the performance and quality of OCR output57
hkust-nlp/cevalAn evaluation suite providing multiple-choice questions for foundation models in various disciplines, with tools for assessing model performance.1,650
krrishdholakia/betterpromptAn API for evaluating the quality of text prompts used in Large Language Models (LLMs) based on perplexity estimation43