poio-corpus

Language dataset

A collection of language resources extracted from publicly available sources.

The Poio Corpus is a freely available collection of language resources for the lesser-used languages. The data is extracted from free sources like Wikipedia, dictionaries, documents, websites and others.

GitHub

7 stars
7 watching
1 forks
Language: Python
last commit: almost 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
cidles/poio-analyzerA collection of software tools for linguists to manage and analyze linguistic data13
cidles/poio-apiA Python library for converting linguistic data from various formats into unified annotation graphs.18
fido-ai/ua-datasetsProvides a collection of datasets for natural language processing in Ukrainian.57
alexa/massiveA collection of tools and modeling code for a large multilingual Natural Language Understanding dataset541
proycon/python-frogA Python binding to a C++ NLP tool for Dutch language processing tasks47
rodrigopivi/chatitoA tool for generating datasets for AI chatbots and natural language processing tasks using a simple domain-specific language.877
dativebase/oldSoftware for creating collaborative databases of language data1
alvations/seedlingA corpus and API for human language data11
01-ai/yiA series of large language models trained from scratch to excel in multiple NLP tasks7,743
clio-lang/clioA functional programming language that compiles to JavaScript and is designed for distributed scientific computing.938
thu-coai/cdial-gptA large-scale Chinese conversation dataset and pre-trained dialog models for text generation1,799
mirfan899/urduA collection of Urdu language datasets for various NLP tasks and applications71
poio-nlp/pressagioA Python library that uses n-gram models to predict text completions19
louisowen6/nlp_bahasa_resourcesA curated collection of NLP datasets and resources for Bahasa Indonesia496
karthikncode/nlp-datasetsA curated list of Natural Language Processing datasets used to train and evaluate NLP models.919