TALPCo

Asian language dataset

A parallel corpus of Asian languages with linguistic annotations and data formats for natural language processing research.

TUFS Asian Language Parallel Corpus

GitHub

49 stars
2 watching
13 forks
Language: TeX
last commit: over 3 years ago
addresseebahasa-indonesiabahasa-melayuburmeseconstituency-treeenglishindonesianinterpersonaljapanesejavanesekoreanmalaymeaningmyanmarparallel-corpusthaitiengviettokenized-sentencestreebankvietnamese

Related projects:

RepositoryDescriptionStars
crownpku/small-chinese-corpusA collection of datasets and tools for NLP tasks on Chinese texts, including part-of-speech tagging, named entity recognition, and question answering.529
kata-ai/indosumProvides a benchmark dataset and tools for training text summarization models in the Indonesian language.77
louisowen6/nlp_bahasa_resourcesA curated collection of NLP datasets and resources for Bahasa Indonesia496
carbonz0/alpaca-chinese-datasetA dataset for training and fine-tuning large language models on Chinese text prompts.392
kyubyong/css10A collection of speech datasets for 10 languages to support text-to-speech tasks467
mirfan899/urduA collection of Urdu language datasets for various NLP tasks and applications71
atik-05/bangla_datasets_absaA collection of pre-processed datasets in Bangla language for natural language processing tasks0
qhungngo/evbcorpusA large-scale bilingual corpus collection for language technology and NLP tasks, containing English-Vietnamese translations and bitexts.42
hit-scir/elmoformanylangsProvides pre-trained ELMo representations for multiple languages to improve NLP tasks.1,462
karthikncode/nlp-datasetsA curated list of Natural Language Processing datasets used to train and evaluate NLP models.919
famrashel/idn-tagged-corpusA manually tagged Indonesian language corpus in tab-separated file format88
lantip/baku-tidak-bakuA repository of linguistic data for Indonesian words categorized as either standard or non-standard29
kangfend/bahasaA natural language processing toolkit for the Indonesian language.19
brightmart/xlnet_zhTrains a large Chinese language model on massive data and provides a pre-trained model for downstream tasks230
chatopera/insuranceqa-corpus-zhAn insurance industry conversation corpus with pre-processed data for natural language processing and question answering tasks.1,019