alpaca-chinese-dataset

Chinese prompt dataset

A dataset for training and fine-tuning large language models on Chinese text prompts.

alpaca中文指令微调数据集

GitHub

392 stars
7 watching
25 forks
last commit: over 3 years ago
alpacachatglmllm

Related projects:

RepositoryDescriptionStars
lc1332/chinese-alpaca-loraDevelops and maintains a Chinese language model finetuned on LLaMA, used for text generation and summarization tasks.711
airaria/visual-chinese-llama-alpacaDevelops a multimodal Chinese language model with visual capabilities429
hikariming/chat-dataset-baselineProvides a resource library for training Chinese conversation models with pre-processed datasets and a framework for fine-tuning the models1,162
icip-cas/chatalpacaA dataset of multi-turn conversations between users and AI models.164
brightmart/xlnet_zhTrains a large Chinese language model on massive data and provides a pre-trained model for downstream tasks230
gururise/alpacadatacleanedA cleaned and curated version of an Alpaca dataset used to train a large language model1,525
pointnetwork/point-alpacaRecreated weights from Stanford Alpaca model fine-tuned for specific task406
crownpku/small-chinese-corpusA collection of datasets and tools for NLP tasks on Chinese texts, including part-of-speech tagging, named entity recognition, and question answering.529
matbahasa/talpcoA parallel corpus of Asian languages with linguistic annotations and data formats for natural language processing research.49
ntunlplab/traditional-chinese-alpacaA research project that develops a Traditional-Chinese instruction-following language model using Alpaca as a basis.134
aisegmentcn/matting_human_datasetsA large dataset of human matting images and corresponding results for training person segmentation models.615
km1994/llmsninestorydemontowerExploring various LLMs and their applications in natural language processing and related areas1,854
cluebenchmark/electraTrains and evaluates a Chinese language model using adversarial training on a large corpus.140
hit-scir/chinese-mixtral-8x7bAn implementation of a large language model for Chinese text processing, focusing on MoE (Multi-Headed Attention) architecture and incorporating a vast vocabulary.645
cluebenchmark/cluecorpus2020A large-scale Chinese corpus for pre-training language models.927