bigcode-evaluation-harness

Code evaluation framework

A framework for evaluating autoregressive code generation language models in terms of their accuracy and robustness.

A framework for the evaluation of autoregressive code generation language models.

GitHub

846 stars
12 watching
225 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
bigcode-project/starcoder2Trains models to generate code in multiple programming languages1,808
modelscope/evalscopeA framework for efficiently evaluating and benchmarking large models308
flageval-baai/flagevalAn evaluation toolkit and platform for assessing large models in various domains307
princeton-nlp/intercodeAn interactive code environment framework for evaluating language agents through execution feedback.198
bin123apple/autocoderAn AI model designed to generate and execute code automatically816
codefuse-ai/codefuse-devops-evalAn evaluation suite for assessing foundation models in the DevOps field.690
relari-ai/continuous-evalProvides a comprehensive framework for evaluating Large Language Model (LLM) applications and pipelines with customizable metrics455
open-evals/evalsA framework for evaluating OpenAI models and an open-source registry of benchmarks.19
quantifiedcode/quantifiedcodeA code analysis and automation platform111
ukgovernmentbeis/inspect_aiA framework for evaluating large language models669
allenai/olmo-evalA framework for evaluating language models on NLP tasks326
quantifiedcode/python-anti-patternsA collection of common Python coding mistakes and poor practices1,716
huggingface/evaluateAn evaluation framework for machine learning models and datasets, providing standardized metrics and tools for comparing model performance.2,063
nvlabs/verilog-evalAn evaluation harness for generating Verilog code from natural language prompts188
budecosystem/code-millenialsA state-of-the-art open-source code generation model with human evaluability score comparable to GPT-4 and Google's proprietary models.20