evals
by openai
Pythonpushed almost 2 years ago
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
AI summary
Benchmarking framework
A framework for evaluating large language models and systems, providing a registry of benchmarks.
- stars
- 15.2K
- forks
- 2.6K
- watching
- 264
- awesome lists
- 4
Featured in 4 awesome lists
Each link jumps to the spot where the list mentions evals.