XiezhiBenchmark
by MikeGu721
AI summary
Questionnaire
An evaluation suite to assess language models' performance in multi-choice questions
- stars
- 93
- forks
- 4
- watching
- 1
Similar projects
Found by comparing what the projects do, not just their names.
Chinese understanding benchmark
Measures the understanding of massive multitask Chinese datasets using large language models
Multilingual Model
Develops and publishes large multilingual language models with advanced mixing-of-experts architecture.
Model evaluator
Evaluates the capabilities of large multimodal models using a set of diverse tasks and metrics
Language Model
A large language model developed to support multiple languages and applications
Language Model
A high-performance language model designed to excel in tasks like natural language understanding, mathematical computation, and code generation
Language Model
A large language model developed by XVERSE Technology Inc. using transformer architecture and fine-tuned on diverse data sets for various applications.
Multimodal model
A large multimodal model for visual question answering, trained on a dataset of 2.1B image-text pairs and 8.2M instruction sequences.
Text classifier
A Python package implementing an interpretable machine learning model for text classification with visualization tools
Video benchmarking toolkit
Evaluates and benchmarks large language models' video understanding capabilities
Medical NLP training data
A large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain
Mixture-of-Experts Model
Developed by XVERSE Technology Inc. as a multilingual large language model with a unique mixture-of-experts architecture and fine-tuned for various tasks such as conversation, question answering, and natural language understanding.
Multimodal model developer
Develops large multimodal models for high-resolution understanding and analysis of text, images, and other data types.
Linguistic analysis toolkit
An open-source framework providing tools and models for analyzing and generating Chinese classics texts using large language models
Multimodal evaluation framework
Develops a multimodal task and dataset to assess vision-language models' ability to handle interleaved image-text inputs.
Model evaluation tool
An analysis project investigating limitations of visual language models in understanding and processing images with potential biases and interference challenges.