EvalAI

Benchmarking tool

A platform for comparing and evaluating AI and machine learning algorithms at scale

cloud rocket bar_chart chart_with_upwards_trend Evaluating state of the art in AI

GitHub

2k stars
54 watching
799 forks
Language: Python
last commit: about 2 years ago
Linked from 1 awesome list

aiai-challengesangular7angularjsartificial-intelligencechallengedjangodockerevalaievaluationleaderboardmachine-learningpythonreproducibilityreproducible-research

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
catboost/benchmarksComparative benchmarks of various machine learning algorithms169
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322
aws-samples/foundation-model-benchmarking-toolA tool for benchmarking performance and accuracy of generative AI models on various AWS platforms210
alco/benchfellaTools for comparing and benchmarking small code snippets514
ethicalml/xaiAn eXplainability toolbox for machine learning that enables data analysis and model evaluation to mitigate biases and improve performance1,135
mshukor/evalign-iclEvaluating and improving large multimodal models through in-context learning21
princeton-nlp/charxivAn evaluation suite for assessing chart understanding in multimodal large language models.85
ys-zong/vl-iclA benchmarking suite for multimodal in-context learning models31
bencheeorg/bencheeA tool for benchmarking Elixir code and comparing performance statistics1,422
bailool/doyouevenlearnA comprehensive resource guide to stay updated on AI, ML, DL, and CV advancements1,039
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84
vlall/swift-brainA collection of algorithms and data structures for artificial intelligence and machine learning in Swift335
openai/simple-evalsEvaluates language models using standardized benchmarks and prompting techniques.2,059
jvalegre/robertAutomated machine learning protocols for cheminformatics using Python39
vchitect/vbenchA benchmark suite for evaluating the performance of video generative models643