AlpacaDataCleaned

Language data set

A cleaned and curated version of an Alpaca dataset used to train a large language model

Alpaca dataset from Stanford, cleaned and curated

GitHub

2k stars
27 watching
153 forks
Language: Python
last commit: over 3 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
alvations/seedlingA corpus and API for human language data11
carbonz0/alpaca-chinese-datasetA dataset for training and fine-tuning large language models on Chinese text prompts.392
pointnetwork/point-alpacaRecreated weights from Stanford Alpaca model fine-tuned for specific task406
alvations/sugarlikeA tool that identifies languages in text by comparing them to a reference set of patterns.1
alvations/sugaliA system designed to identify the language of an arbitrary text string using machine learning and multiple data sources.2
google-research/flanA repository providing tools and datasets to fine-tune language models for specific tasks1,484
code-kern-ai/refineryA tool to help data scientists manage and annotate natural language data for training AI models1,405
flagai-open/aquila2Provides pre-trained language models and tools for fine-tuning and evaluation439
matbahasa/talpcoA parallel corpus of Asian languages with linguistic annotations and data formats for natural language processing research.49
datacanvasio/alayaA pre-trained AI model that can engage in natural language conversations with high accuracy and understanding.43
airaria/visual-chinese-llama-alpacaDevelops a multimodal Chinese language model with visual capabilities429
karthikncode/nlp-datasetsA curated list of Natural Language Processing datasets used to train and evaluate NLP models.919
alpacahq/alpaca-trade-api-pythonA Python client for Alpaca's trade API1,745
vhellendoorn/code-lmsA guide to using pre-trained large language models in source code analysis and generation1,789
sparklingpandas/sparklingpandasEnables distributed data analysis using PySpark and Pandas APIs362