GAOKAO-Bench
by OpenLMLab
GAOKAO-Bench is an evaluation framework that utilizes GAOKAO questions as a dataset to evaluate large language models.
AI summary
Model testing framework
An evaluation framework using Chinese high school examination questions to assess large language model capabilities
- stars
- 565
- forks
- 42
- watching
- 4
Similar projects
Found by comparing what the projects do, not just their names.
Chinese understanding benchmark
Measures the understanding of massive multitask Chinese datasets using large language models
Language model toolkit
Provides pre-trained language models and tools for fine-tuning and evaluation
Model evaluator
Evaluates and compares the performance of multimodal large language models on various tasks
zjunlp/knowlm1.3K
Knowledge model framework
A framework for training and utilizing large language models with knowledge augmentation capabilities
AI agent framework
A framework and benchmark for training and evaluating multi-modal large language models, enabling the development of AI agents capable of seamless interaction between humans and machines.
Visual Model Benchmark
An open-source benchmarking framework for evaluating cross-style visual capability of large multimodal models
LLaMA patch
An incremental pre-trained Chinese large language model based on the LLaMA-7B model
Test framework
A testing framework that enables the creation of long-running tests using a domain-specific language.
Legal model evaluator
Evaluates the legal knowledge of large language models using a custom benchmarking framework.
Testing Framework
A cross-platform desktop testing framework utilizing accessibility APIs and the JNA library to automate interactions with GUI elements.
google/paxml461
ML framework
A framework for configuring and running machine learning experiments on top of Jax.
Model Tester
A tool for testing and evaluating large language models with a focus on AI safety and model assessment.
Video benchmarking toolkit
Evaluates and benchmarks large language models' video understanding capabilities
Image captioner
An end-to-end image captioning system that uses large multi-modal models and provides tools for training, inference, and demo usage.
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics