CLUECorpus2020
Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料
AI summary
Corpus
A large-scale Chinese corpus for pre-training language models.
- stars
- 927
- forks
- 81
- watching
- 20
Similar projects
Found by comparing what the projects do, not just their names.
Chinese language models
Provides pre-trained models for Chinese language tasks with improved performance and smaller model sizes compared to existing models.
Chinese Language Model
Trains and evaluates a Chinese language model using adversarial training on a large corpus.
NLP Model
A pre-trained language model for multiple natural language processing tasks with support for few-shot learning and transfer learning.
Model benchmark
A benchmarking platform for evaluating Chinese general-purpose models through anonymous, random battles
Chinese language model
Trains a large Chinese language model on massive data and provides a pre-trained model for downstream tasks
clue-ai/chatyuan1.9K
Language dialog model
Large language model for dialogue support in multiple languages
Chinese text dataset suite
A collection of datasets and tools for NLP tasks on Chinese texts, including part-of-speech tagging, named entity recognition, and question answering.
NLP multi-task dataset
A large-scale dataset for training models to perform multiple tasks and zero-shot learning in natural language processing.
Language Model Upgrade
An updated version of a large language model designed to improve performance on multiple tasks and datasets
Chinese character understanding model
A deep learning model that incorporates visual and phonetic features of Chinese characters to improve its ability to understand Chinese language nuances
News corpus
A large dataset of news articles with labeled categories to train fake news recognition algorithms
Chinese character recognition network
This project demonstrates how to build and train a convolutional neural network (CNN) to recognize Chinese characters.
Chinese word embedding trainer
This is a software project that trains and evaluates word embeddings for Chinese words, characters, and fine-grained subcharacter components.
Chinese language model fine-tuning tool
Improves pre-trained Chinese language models by incorporating a correction task to alleviate inconsistency issues with downstream tasks
Forum dataset
A collection of question-answer pairs extracted from online Chinese forums.