chazutsu

NLP dataset manager

A tool that simplifies the process of preparing and manipulating natural language processing datasets

The tool to make NLP datasets ready to use

GitHub

243 stars
14 watching
33 forks
Language: Python
last commit: almost 4 years ago
Linked from 1 awesome list

datasetmachine-learningnatural-language-processing

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
karthikncode/nlp-datasetsA curated list of Natural Language Processing datasets used to train and evaluate NLP models.919
kimtaro/veA linguistic framework for natural language processing tasks.216
chartbeat-labs/textacyA Python library providing NLP tools and utilities built on top of spaCy for text processing and analysis.2,217
alexa/massiveA collection of tools and modeling code for a large multilingual Natural Language Understanding dataset541
languagemachines/luiginlpA workflow management system for Natural Language Processing tasks21
goru001/inltkA comprehensive toolkit for Natural Language Processing tasks in Indic languages, providing pre-trained models and datasets.825
jd-aig/nlp_baaiA collection of natural language processing models and tools for collaboration on a joint project between BAAI and JDAI.254
radi-cho/datasetgptA command-line interface to generate textual datasets with Large Language Models293
pks/zipfA Ruby NLP library providing tools and data structures for natural language processing tasks3
michael-wzhu/promptcblueA large-scale instruction-tuning dataset for multi-task and few-shot learning in the medical domain328
piskvorky/gensim-dataA repository of pre-trained NLP models and corpora for text processing.990
dayyass/dayyassA collection of libraries and tools for natural language processing and reinforcement learning.39
dkogan/vnlogA toolkit for manipulating tabular ASCII data with normal UNIX tools.161
zaibacu/rita-dslA DSL for building custom NLP patterns from manual language rules65
leks-forever/nllb-tuningThis is an experimental project for fine-tuning the NLB language model with a specific dataset and evaluating its performance on translation tasks.7