auto-evaluator

QA evaluation tool

Automated evaluation of language models for question answering tasks

GitHub

749 stars
12 watching
100 forks
Language: TypeScript
last commit: over 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
rlancemartin/auto-evaluatorAn evaluation tool for question-answering systems using large language models and natural language processing techniques1,065
allenai/document-qaTools and codebase for training neural question answering models on multiple paragraphs of text data435
retraigo/appraisalUtilities for transforming and analyzing text data using machine learning algorithms5
langchain-ai/langserveProvides a REST API for deploying and managing LangChain runnables and chains1,970
mlabonne/llm-autoevalA tool to automate the evaluation of large language models in Google Colab using various benchmarks and custom parameters.566
allenai/olmo-evalA framework for evaluating language models on NLP tasks326
cloud-cv/evalaiA platform for comparing and evaluating AI and machine learning algorithms at scale1,779
open-compass/lawbenchEvaluates the legal knowledge of large language models using a custom benchmarking framework.273
ruixiangcui/agievalEvaluates foundation models on human-centric tasks with diverse exams and question types714
wordweb/langchain-chatglm-and-tigerbotDevelops a knowledge-based question answering application using Open Source models like ChatGLM and TigerBot105
kevincoble/aitoolboxA toolbox of AI modules written in Swift for various machine learning tasks and algorithms794
openai/simple-evalsEvaluates language models using standardized benchmarks and prompting techniques.2,059
langchain-ai/langgraphjsA framework for building resilient, stateful applications with LLMs as directed graphs742
johnsnowlabs/langtestA tool for testing and evaluating large language models with a focus on AI safety and model assessment.506
tatsu-lab/alpaca_evalAn automatic evaluation tool for large language models1,568