BIG-bench

Language model benchmark

A benchmark designed to probe large language models and extrapolate their future capabilities through a diverse set of tasks.

Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models

GitHub

3k stars
51 watching
593 forks
Language: Python
last commit: about 2 years ago
Linked from 2 awesome lists


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
bigscience-workshop/promptsourceA toolkit for creating and using natural language prompts to enable large language models to generalize to new tasks.2,718
brexhq/prompt-engineeringGuides software developers on how to effectively use and build systems around Large Language Models like GPT-4.8,487
kostya/benchmarksA collection of benchmarking tests for various programming languages2,825
fminference/flexllmgenGenerates large language model outputs in high-throughput mode on single GPUs9,236
microsoft/promptbenchA unified framework for evaluating large language models' performance and robustness in various scenarios.2,487
openbmb/bmtoolsTools and platform for building and extending large language models2,907
huggingface/text-generation-inferenceA toolkit for deploying and serving Large Language Models (LLMs) for high-performance text generation9,456
google/benchmarkA microbenchmarking library that allows users to measure the execution time of specific code snippets9,113
optimalscale/lmflowA toolkit for fine-tuning and inferring large machine learning models8,312
openbmb/toolbenchA platform for training, serving, and evaluating large language models to enable tool use capability4,888
brightmart/text_classificationAn NLP project offering various text classification models and techniques for deep learning exploration7,881
felixgithub2017/mmcuMeasures the understanding of massive multitask Chinese datasets using large language models87
deepseek-ai/deepseek-v2A high-performance mixture-of-experts language model with strong performance and efficient inference capabilities.3,758
confident-ai/deepevalA framework for evaluating large language models4,003
tianyi-lab/hallusionbenchAn image-context reasoning benchmark designed to challenge large vision-language models and help improve their accuracy259