M3Exam

LM Benchmark

A benchmark for evaluating large language models in multiple languages and formats

Data and code for paper "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models"

GitHub

93 stars
9 watching
12 forks
Language: Python
last commit: over 3 years ago
ai-educationchatgptevaluationgpt-4large-language-modelsllmsmultilingualmultimodal

Related projects:

RepositoryDescriptionStars
damo-nlp-mt/polylmA polyglot large language model designed to address limitations in current LLM research and provide better multilingual instruction-following capability.77
damo-nlp-sg/llm-zooA collection of information about various large language models used in natural language processing272
qcri/llmebenchA benchmarking framework for large language models81
aifeg/benchlmmAn open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models84
pleisto/yuren-baichuan-7bA multi-modal large language model that integrates natural language and visual capabilities with fine-tuning for various tasks73
junyangwang0410/amberAn LLM-free benchmark suite for evaluating MLLMs' hallucination capabilities in various tasks and dimensions98
ray-project/llmperfA tool for evaluating the performance of large language model APIs678
mlgroupjlu/llm-eval-surveyA repository of papers and resources for evaluating large language models.1,450
deeplangai/lingowhale-8bAn open bilingual LLM developed using the LingoWhale model, trained on a large dataset of high-quality middle English text, and fine-tuned for specific tasks such as conversation generation.134
km1994/llmsninestorydemontowerExploring various LLMs and their applications in natural language processing and related areas1,854
bobazooba/xllmA tool for training and fine-tuning large language models using advanced techniques387
bilibili/index-1.9bA lightweight, multilingual language model with a long context length920
damoebius/haxebenchA benchmarking project comparing the performance of different programming languages and their compiled outputs in various formats.52
deepseek-ai/deepseek-moeA large language model with improved efficiency and performance compared to similar models1,024
ailab-cvc/seed-benchA benchmark for evaluating large language models' ability to process multimodal input322