codefuse-devops-eval

DevOps benchmark

An evaluation suite for assessing foundation models in the DevOps field.

Industrial-first evaluation benchmark for LLMs in the DevOps/AIOps domain.

GitHub

690 stars
9 watching
44 forks
Language: Python
last commit: about 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
codefuse-ai/codefuse-devops-modelAn industrial-first language model for answering questions in the DevOps domain596
codefuse-ai/codefuse-chatbotAn AI-powered tool designed to simplify and optimize various stages of the software development lifecycle1,202
codefuse-ai/test-agentA tool that empowers software testing with large language models.565
codefuse-ai/mftcoderA framework for fine-tuning large language models with multiple tasks to improve their accuracy and efficiency647
bregman-arie/howtheydevopsA collection of publicly available resources on DevOps practices from companies around the world733
princeton-nlp/intercodeAn interactive code environment framework for evaluating language agents through execution feedback.198
open-evals/evalsA framework for evaluating OpenAI models and an open-source registry of benchmarks.19
cloud-cv/evalaiA platform for comparing and evaluating AI and machine learning algorithms at scale1,779
hkust-nlp/cevalAn evaluation suite providing multiple-choice questions for foundation models in various disciplines, with tools for assessing model performance.1,650
johnathan79717/codeforces-parserGenerates sample tests and input/output files for competitive programming contests137
alco/benchfellaTools for comparing and benchmarking small code snippets514
microsoft/codexglueA benchmark dataset and open challenge to improve AI models' ability to understand and generate code1,575
openai/simple-evalsEvaluates language models using standardized benchmarks and prompting techniques.2,059
joelwmale/codeception-actionAn action for running Codeception tests in GitHub workflows15
openai/procgenA benchmark for evaluating reinforcement learning agent performance on procedurally generated game-like environments.1,030