SuperCLUElyb

Model benchmark

A benchmarking platform for evaluating Chinese general-purpose models through anonymous, random battles

SuperCLUE琅琊榜:中文通用大模型匿名对战评价基准

GitHub

143 stars
5 watching
6 forks
last commit: over 2 years ago

Related projects:

RepositoryDescriptionStars
cluebenchmark/cluepretrainedmodelsProvides pre-trained models for Chinese language tasks with improved performance and smaller model sizes compared to existing models.806
cluebenchmark/cluecorpus2020A large-scale Chinese corpus for pre-training language models.927
cluebenchmark/electraTrains and evaluates a Chinese language model using adversarial training on a large corpus.140
cluebenchmark/pclueA large-scale dataset for training models to perform multiple tasks and zero-shot learning in natural language processing.473
clue-ai/promptclueA pre-trained language model for multiple natural language processing tasks with support for few-shot learning and transfer learning.656
felixgithub2017/mmcuMeasures the understanding of massive multitask Chinese datasets using large language models87
clue-ai/chatyuan-7bAn updated version of a large language model designed to improve performance on multiple tasks and datasets13
qcri/llmebenchA benchmarking framework for large language models81
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84
catboost/benchmarksComparative benchmarks of various machine learning algorithms169
bitshifter/mathbench-rsA benchmarking framework comparing performance of different Rust linear algebra libraries200
ibob/picobenchA microbenchmarking library for C++211
yuliang-liu/multimodalocrAn evaluation benchmark for OCR capabilities in large multmodal models.484
robustbench/robustbenchA standardized benchmark for measuring the robustness of machine learning models against adversarial attacks682
clue-ai/chatyuanLarge language model for dialogue support in multiple languages1,903