wongnai-corpus

Thai NLP Datasets

A collection of datasets for natural language processing research in Thai, including word segmentation and review rating prediction.

Collection of Wongnai's datasets

GitHub

76 stars
6 watching
23 forks
last commit: about 7 years ago
Linked from 1 awesome list

datasetsnlpnlp-machine-learningtokenization

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
louisowen6/nlp_bahasa_resourcesA curated collection of NLP datasets and resources for Bahasa Indonesia496
karthikncode/nlp-datasetsA curated list of Natural Language Processing datasets used to train and evaluate NLP models.919
krakenai/synthaiA deep learning-based project for segmenting Thai text into words and annotating parts of speech with high accuracy.41
pythainlp/lexicon-thaiA Thai language corpus and lexicon repository for natural language processing142
mirfan899/urduA collection of Urdu language datasets for various NLP tasks and applications71
pythainlp/pythainlpA Python package for text processing and linguistic analysis focused on Thai language993
tmu-nlp/thaitoxicitytweetcorpusCorpus of annotated Thai tweets to analyze toxicity and sentiment10
vinairesearch/phobertPre-trained language models for Vietnamese NLP tasks671
crownpku/small-chinese-corpusA collection of datasets and tools for NLP tasks on Chinese texts, including part-of-speech tagging, named entity recognition, and question answering.529
wannaphong/thai-nerNamed Entity Recognition for Thai Text using PyThaiNLP and custom implementation.53
pythainlp/prachathai-67kAn article classification dataset created from news articles scraped from Prachathai.com with multiple benchmark models for multi-label classification16
rkcosmos/deepcutA Thai word tokenization library using Deep Neural Network421
ymcui/chinese-xlnetProvides pre-trained models for Chinese natural language processing tasks using the XLNet architecture1,652
matbahasa/talpcoA parallel corpus of Asian languages with linguistic annotations and data formats for natural language processing research.49
zhuiyitechnology/pretrained-modelsA collection of pre-trained language models for natural language processing tasks989