h2o-LLM-eval
by h2oai
Large-language Model Evaluation framework with Elo Leaderboard and A-B testing
AI summary
LLM evaluator
An evaluation framework for large language models with Elo rating system and A/B testing capabilities
- stars
- 50
- forks
- 1
- watching
- 37
Similar projects
Found by comparing what the projects do, not just their names.
Evaluation pipeline
A framework for evaluating language models on NLP tasks
LLM evaluation resource
A repository of papers and resources for evaluating large language models.
Model interpreter
Provides tools and techniques for interpreting machine learning models
LLM evaluation framework
Provides a comprehensive framework for evaluating Large Language Model (LLM) applications and pipelines with customizable metrics
ML platform
An in-memory machine learning platform that supports various algorithms and provides tools for building, deploying, and scaling machine learning models
Model evaluation toolkit
Tools and evaluation framework for accelerating the development of large multimodal models by providing an efficient way to assess their performance
ML framework
A framework for building and evaluating machine learning systems with high accuracy and interpretability, particularly in human-centered applications.
Model evaluator
An evaluation framework for large language models trained with instruction tuning methods
LLM evaluator collection
A collaborative catalogue of LLM evaluation frameworks and papers
h2oai/h2o-22.2K
Analytics Engine
An analytics engine that provides fast and scalable predictive modeling capabilities for big data
Model evaluator
A tool to automate the evaluation of large language models in Google Colab using various benchmarks and custom parameters.
Evaluator
An automatic evaluation tool for large language models
Computing environment
An interactive computing environment for machine learning and data analysis
LLM evaluator
An all-in-one toolkit for evaluating Large Language Models (LLMs) across multiple backends.
Model Evaluator
A benchmarking framework for evaluating Large Multimodal Models by providing rigorous metrics and an efficient evaluation pipeline.