Awesome Lists

awesome-nlp

by keon

awesome listpushed almost 3 years ago

book A curated list of resources dedicated to Natural Language Processing (NLP)

AI summary

NLP resource repository

A curated collection of resources and references for Natural Language Processing

stars
16.8K
forks
2.6K
watching
609
awesome lists
8
entries
326
View on GitHub

Embed the badge

Show how many awesome lists link to your project. The count updates automatically.

Awesome Lists badge
Markdown
[![Awesome Lists Badge](https://awesome.facts.dev/shield/keon/awesome-nlp/links.svg)](https://awesome.facts.dev/awesome/keon/awesome-nlp)
HTML
<a href="https://awesome.facts.dev/awesome/keon/awesome-nlp"><img src="https://awesome.facts.dev/shield/keon/awesome-nlp/links.svg" alt="Awesome Lists Badge" /></a>
Image URL
https://awesome.facts.dev/shield/keon/awesome-nlp/links.svg

What's in the list

326 links in 47 sections, with live GitHub stats.activeno commit in 2y

Prominent NLP Research Labs

Tutorials / Reading Content

Tutorials / Videos and Online Courses

Tutorials / Books

Libraries

  • Twitter-text

    A JavaScript implementation of Twitter's text processing library

  • Knwl.js

    A Natural Language Processor in JS

  • Retext

    Extensible system for analyzing and manipulating natural language

  • NLP Compromise

    Natural Language processing in the browser

  • Natural

    general natural language facilities for node

  • Poplar

    A web-based annotation tool for natural language processing (NLP)

  • NLP.js

    An NLP library for building bots

  • node-question-answering

    Fast and production-ready question answering w/ DistilBERT in Node.js

  • sentimental-onix

    Sentiment models for spacy using onnx

  • TextAttack

    Adversarial attacks, adversarial training, and data augmentation in NLP

  • TextBlob

    Providing a consistent API for diving into common natural language processing (NLP) tasks. Stands on the giant shoulders of and , and plays nicely with both

  • spaCy

    Industrial strength NLP with Python and Cython

  • Speedster

    Automatically apply SOTA optimization techniques to achieve the maximum inference speed-up on your hardware

Libraries / Speedster

  • textacy

    Higher level NLP built on spaCy

Libraries

  • gensim

    Python library to conduct unsupervised semantic modelling from plain text

  • scattertext

    Python library to produce d3 visualizations of how language differs between corpora

  • GluonNLP

    A deep learning toolkit for NLP, built on MXNet/Gluon, for research prototyping and industrial deployment of state-of-the-art models on a wide range of NLP tasks

  • AllenNLP

    An NLP research library, built on PyTorch, for developing state-of-the-art deep learning models on a wide variety of linguistic tasks

  • PyTorch-NLP

    NLP research toolkit designed to support rapid prototyping with better data loaders, word vector loaders, neural network layer representations, common NLP metrics such as BLEU

  • Rosetta

    Text processing tools and wrappers (e.g. Vowpal Wabbit)

  • PyNLPl

    Python Natural Language Processing Library. General purpose NLP library for Python, handles some specific formats like ARPA language models, Moses phrasetables, GIZA++ alignments

  • foliapy

    Python library for working with , an XML format for linguistic annotation

  • PySS3

    Python package that implements a novel white-box machine learning model for text classification, called SS3. Since SS3 has the ability to visually explain its rationale, this package also comes with easy-to-use interactive visualizations tools ( )

  • jPTDP

    A toolkit for joint part-of-speech (POS) tagging and dependency parsing. jPTDP provides pre-trained models for 40+ languages

  • BigARTM

    a fast library for topic modelling

  • Snips NLU

    A production ready library for intent parsing

  • Chazutsu

    A library for downloading&parsing standard NLP research datasets

  • Word Forms

    Word forms can accurately generate all possible forms of an English word

  • Multilingual Latent Dirichlet Allocation (LDA)

    A multilingual and extensible document clustering pipeline

  • Natural Language Toolkit (NLTK)

    A library containing a wide variety of NLP functionality, supporting over 50 corpora

  • NLP Architect

    A library for exploring the state-of-the-art deep learning topologies and techniques for NLP and NLU

  • Flair

    A very simple framework for state-of-the-art multilingual NLP built on PyTorch. Includes BERT, ELMo and Flair embeddings

  • Kashgari

    Simple, Keras-powered multilingual NLP framework, allows you to build your models in 5 minutes for named entity recognition (NER), part-of-speech tagging (PoS) and text classification tasks. Includes BERT and word2vec embedding

  • FARM

    Fast & easy transfer learning for NLP. Harvesting language models for the industry. Focus on Question Answering

  • Haystack

    End-to-end Python framework for building natural language search interfaces to data. Leverages Transformers and the State-of-the-Art of NLP. Supports DPR, Elasticsearch, HuggingFace’s Modelhub, and much more!

  • Rita DSL

    a DSL, loosely based on . Allows to define language patterns (rule-based NLP) which are then translated into , or if you prefer less features and lightweight - regex patterns

  • Transformers

    Natural Language Processing for TensorFlow 2.0 and PyTorch

  • Tokenizers

    Tokenizers optimized for Research and Production

  • fairSeq

    Facebook AI Research implementations of SOTA seq2seq models in Pytorch

  • corex_topic

    Hierarchical Topic Modeling with Minimal Domain Knowledge

  • Sockeye

    Neural Machine Translation (NMT) toolkit that powers Amazon Translate

  • DL Translate

    A deep learning-based translation library for 50 languages, built on and Facebook's mBART Large

  • Jury

    Evaluation of NLP model outputs offering various automated metrics

  • python-ucto

    Unicode-aware regular-expression based tokenizer for various languages. Python binding to C++ library, supports

  • InsNet

    A neural network library for building instance-dependent NLP models with padding-free dynamic batching

  • MIT Information Extraction Toolkit

    C, C++, and Python tools for named entity recognition and relation extraction

  • CRF++

    Open source implementation of Conditional Random Fields (CRFs) for segmenting/labeling sequential data & other Natural Language Processing tasks

  • CRFsuite

    CRFsuite is an implementation of Conditional Random Fields (CRFs) for labeling sequential data

  • BLLIP Parser

    BLLIP Natural Language Parser (also known as the Charniak-Johnson parser)

  • colibri-core

    C++ library, command line tools, and Python binding for extracting and working with basic linguistic constructions such as n-grams and skipgrams in a quick and memory-efficient way

  • ucto

    Unicode-aware regular-expression based tokenizer for various languages. Tool and C++ library. Supports FoLiA format

  • libfolia

    C++ library for the

  • frog

    Memory-based NLP suite developed for Dutch: PoS tagger, lemmatiser, dependency parser, NER, shallow parser, morphological analyzer

  • MeTA

    is a C++ Data Sciences Toolkit that facilitates mining big text data

  • StarSpace

    a library from Facebook for creating embeddings of word-level, paragraph-level, document-level and for text classification

  • ReVerb

    Web-Scale Open Information Extraction

  • OpenRegex

    An efficient and flexible token-based regular expression language and engine

  • CogcompNLP

    Core libraries developed in the U of Illinois' Cognitive Computation Group

  • MALLET

    MAchine Learning for LanguagE Toolkit - package for statistical natural language processing, document classification, clustering, topic modeling, information extraction, and other machine learning applications to text

  • RDRPOSTagger

    A robust POS tagging toolkit available (in both Java & Python) together with pre-trained models for 40+ languages

  • Lingua

    A language detection library for Kotlin and Java, suitable for long and short text alike

  • Kotidgy

    — an index-based text data generator written in Kotlin

  • Saul

    Library for developing NLP systems, including built in modules like SRL, POS, etc

  • ATR4S

    Toolkit with state-of-the-art methods

  • tm

    Implementation of topic modeling based on regularized multilingual

  • word2vec-scala

    Scala interface to word2vec model; includes operations on vectors like word-distance and word-analogy

  • Epic

    Epic is a high performance statistical parser written in Scala, along with a framework for building complex structured prediction models

  • Spark NLP

    Spark NLP is a natural language processing library built on top of Apache Spark ML that provides simple, performant & accurate NLP annotations for machine learning pipelines that scale easily in a distributed environment

  • text2vec

    Fast vectorization, topic modeling, distances and GloVe word embeddings in R

  • wordVectors

    An R package for creating and exploring word2vec and other word embedding models

  • RMallet

    R package to interface with the Java machine learning tool MALLET

  • dfr-browser

    Creates d3 visualizations for browsing topic models of text in a web browser

  • dfrtopics

    R package for exploring topic models of text

  • sentiment_classifier

    Sentiment Classification using Word Sense Disambiguation and WordNet Reader

  • jProcessing

    Japanese Natural Langauge Processing Libraries, with Japanese sentiment classification

  • corporaexplorer

    An R package for dynamic exploration of text collections

  • tidytext

    Text mining using tidy tools

  • spacyr

    R wrapper to spaCy NLP

  • Clojure-openNLP

    Natural Language Processing in Clojure (opennlp)

  • Infections-clj

    Rails-like inflection library for Clojure and ClojureScript

  • postagga

    A library to parse natural language in Clojure and ClojureScript

  • whatlang

    — Natural language recognition library based on trigrams

  • snips-nlu-rs

    A production ready library for intent parsing

  • rust-bert

    Ready-to-use NLP pipelines and Transformer-based models

  • VSCode Language Extension

    NLP++ Language Extension for VSCode

  • nlp-engine

    NLP++ engine to run NLP++ code on Linux including a full English parser

  • VisualText

    Homepage for the NLP++ Language

  • NLP++ Wiki

    Wiki entry for the NLP++ language

  • CorpusLoaders

    A variety of loaders for various NLP corpora

  • Languages

    A package for working with human languages

  • TextAnalysis

    Julia package for text analysis

  • TextModels

    Neural Network based models for Natural Language Processing

  • WordTokenizers

    High performance tokenizers for natural language processing and other related tasks

  • Word2Vec

    Julia interface to word2vec

Libraries / Services

  • Wit-ai

    Natural Language Interface for apps and devices

  • Amazon Comprehend

    NLP and ML suite covers most common tasks like NER, tagging, and sentiment analysis

  • Google Cloud Natural Language API

    Syntax Analysis, NER, Sentiment Analysis, and Content tagging in atleast 9 languages include English and Chinese (Simplified and Traditional)

  • ParallelDots

    High level Text Analysis API Service ranging from Sentiment Analysis to Intent Analysis

  • Textalytic

    Natural Language Processing in the Browser with sentiment analysis, named entity extraction, POS tagging, word frequencies, topic modeling, word clouds, and more

  • NLP Cloud

    SpaCy NLP models (custom and pre-trained ones) served through a RESTful API for named entity recognition (NER), POS tagging, and more

  • Cloudmersive

    Unified and free NLP APIs that perform actions such as speech tagging, text rephrasing, language translation/detection, and sentence parsing

Libraries / Annotation Tools

  • GATE

    General Architecture and Text Engineering is 15+ years old, free and open source

  • Anafora

    is free and open source, web-based raw text annotation tool

  • brat

    brat rapid annotation tool is an online environment for collaborative text annotation

  • doccano

    doccano is free, open-source, and provides annotation features for text classification, sequence labeling and sequence to sequence

  • INCEpTION

    A semantic annotation platform offering intelligent assistance and knowledge management

  • tagtog

    , team-first web tool to find, create, maintain, and share datasets - costs $

  • prodigy

    is an annotation tool powered by active learning, costs $

  • LightTag

    Hosted and managed text annotation tool for teams, costs $

  • rstWeb

    open source local or online tool for discourse tree annotations

  • GitDox

    open source server annotation tool with GitHub version control and validation for XML data and collaborative spreadsheet grids

  • Label Studio

    Hosted and managed text annotation tool for teams, freemium based, costs $

  • Datasaur

    support various NLP tasks for individual or teams, freemium based

  • Konfuzio

    team-first hosted and on-prem text, image and PDF annotation tool powered by active learning, freemium based, costs $

  • UBIAI

    Easy-to-use text annotation tool for teams with most comprehensive auto-annotation features. Supports NER, relations and document classification as well as OCR annotation for invoice labeling, costs $

  • Shoonya

    Shoonya is free and open source data annotation platform with wide varities of organization and workspace level management system. Shoonya is data agnostic, can be used by teams to annotate data with various level of verification stages at scale

  • Annotation Lab

    Free End-to-End No-Code platform for text annotation and DL model training/tuning. Out-of-the-box support for Named Entity Recognition, Classification, Relation extraction and Assertion Status Spark NLP models. Unlimited support for users, teams, projects, documents. Not FOSS

  • FLAT

    FLAT is a web-based linguistic annotation environment based around the , a rich XML-based format for linguistic annotation. Free and open source

Techniques / Text Embeddings

Techniques / Question Answering and Knowledge Extraction

Datasets

  • nlp-datasets

    great collection of nlp datasets

  • gensim-data

    Data repository for pretrained NLP models and NLP corpora

Multilingual NLP Frameworks

  • UDPipe

    is a trainable pipeline for tokenizing, tagging, lemmatizing and parsing Universal Treebanks and other CoNLL-U files. Primarily written in C++, offers a fast and reliable solution for multilingual NLP processing

  • NLP-Cube

    : Natural Language Processing Pipeline - Sentence Splitting, Tokenization, Lemmatization, Part-of-speech Tagging and Dependency Parsing. New platform, written in Python with Dynet 2.0. Offers standalone (CLI/Python bindings) and server functionality (REST API)

  • UralicNLP

    is an NLP library mostly for many endangered Uralic languages such as Sami languages, Mordvin languages, Mari languages, Komi languages and so on. Also some non-endangered languages are supported such as Finnish together with non-Uralic languages such as Swedish and Arabic. UralicNLP can do morphological analysis, generation, lemmatization and disambiguation

NLP in Korean / Libraries

  • KoNLPy

    Python package for Korean natural language processing

  • Mecab (Korean)

    C++ library for Korean NLP

  • KoalaNLP

    Scala library for Korean Natural Language Processing

  • KoNLP

    R package for Korean Natural language processing

NLP in Korean / Blogs and Tutorials

NLP in Korean / Datasets

NLP in Arabic / Libraries

  • goarabic

    Go package for Arabic text processing

  • jsastem

    Javascript for Arabic stemming

  • PyArabic

    Python libraries for Arabic

  • RFTokenizer

    trainable Python segmenter for Arabic, Hebrew and Coptic

NLP in Arabic / Datasets

  • Multidomain Datasets

    Largest Available Multi-Domain Resources for Arabic Sentiment Analysis

  • LABR

    LArge Arabic Book Reviews dataset

  • Arabic Stopwords

    A list of Arabic stopwords from various resources

NLP in Chinese / Libraries

  • jieba

    Python package for Words Segmentation Utilities in Chinese

  • SnowNLP

    Python package for Chinese NLP

  • FudanNLP

    Java library for Chinese text processing

  • HanLP

    The multilingual NLP library

NLP in Chinese / Anthology

  • funNLP

    Collection of NLP tools and resources mainly for Chinese

NLP in German

  • German-NLP

    Curated list of open-access/open-source/off-the-shelf resources and tools developed with a particular focus on German

NLP in Polish

  • Polish-NLP

    A curated list of resources dedicated to Natural Language Processing (NLP) in polish. Models, tools, datasets

NLP in Spanish / Libraries

  • spanlp

    Python library to detect, censor and clean profanity, vulgarities, hateful words, racism, xenophobia and bullying in texts written in Spanish. It contains data of 21 Spanish-speaking countries

NLP in Spanish / Data

NLP in Spanish / Word and Sentence Embeddings

NLP in Indic languages / Data, Corpora and Treebanks

NLP in Indic languages / Data, Corpora and Treebanks / Universal Dependencies Treebank in Hindi

NLP in Indic languages / Data, Corpora and Treebanks

NLP in Indic languages / Language Models and Word Embeddings

NLP in Indic languages / Libraries and Tooling

  • Multi-Task Deep Morphological Analyzer

    Deep Network based Morphological Parser for Hindi and Urdu

  • Anoop Kunchukuttan

    18 Languages, whole host of features from tokenization to translation

  • SivaReddy's Dependency Parser

    Dependency Parser and Pos Tagger for Kannada, Hindi and Telugu

  • iNLTK

    A Natural Language Toolkit for Indic Languages (Indian subcontinent languages) built on top of Pytorch/Fastai, which aims to provide out of the box support for common NLP tasks

NLP in Thai / Libraries

  • PyThaiNLP

    Thai NLP in Python Package

  • JTCC

    A character cluster library in Java

  • CutKum

    Word segmentation with deep learning in TensorFlow

  • Thai Language Toolkit

    Based on a paper by Wirote Aroonmanakun in 2002 with included dataset

  • SynThai

    Word segmentation and POS tagging using deep learning in Python

NLP in Thai / Data

  • Inter-BEST

    A text corpus with 5 million words with word segmentation

  • Prime Minister 29

    Dataset containing speeches of the current Prime Minister of Thailand

NLP in Danish

NLP in Vietnamese / Libraries

  • underthesea

    Vietnamese NLP Toolkit

  • vn.vitk

    A Vietnamese Text Processing Toolkit

  • VnCoreNLP

    A Vietnamese natural language processing toolkit

  • PhoBERT

    Pre-trained language models for Vietnamese

  • pyvi

    Python Vietnamese Core NLP Toolkit

NLP in Vietnamese / Data

  • Vietnamese treebank

    10,000 sentences for the constituency parsing task

  • BKTreeBank

    a Vietnamese Dependency Treebank

  • UD_Vietnamese

    Vietnamese Universal Dependency Treebank

  • VIVOS

    a free Vietnamese speech corpus consisting of 15 hours of recording speech by AILab

  • VNTQcorpus(big).txt

    1.75 million sentences in news

  • ViText2SQL

    A dataset for Vietnamese Text-to-SQL semantic parsing (EMNLP-2020 Findings)

  • EVB Corpus

    20,000,000 words (20 million) from 15 bilingual books, 100 parallel English-Vietnamese / Vietnamese-English texts, 250 parallel law and ordinance texts, 5,000 news articles, and 2,000 film subtitles

NLP for Dutch

  • python-frog

    Python binding to Frog, an NLP suite for Dutch. (pos tagging, lemmatisation, dependency parsing, NER)

  • SimpleNLG_NL

    Dutch surface realiser used for Natural Language Generation in Dutch, based on the SimpleNLG implementation for English and French

  • Alpino

    Dependency parser for Dutch (also does PoS tagging and Lemmatisation)

  • Kaldi NL

    Dutch Speech Recognition models based on

  • spaCy

    available. - Industrial strength NLP with Python and Cython

NLP in Indonesian / Datasets

  • ILPS

    Kompas and Tempo collections at

  • PANL10N for PoS tagging

    : 39K sentences and 900K word tokens

  • IDN for PoS tagging

    : This corpus contains 10K sentences and 250K word tokens

  • IndoSum

    for text summarization and classification both

  • Wordnet-Bahasa

    large, free, semantic dictionary

  • IndoNLU

    IndoBenchmark includes pre-trained language model (IndoBERT), FastText model, Indo4B corpus, and several NLU benchmark datasets

NLP in Indonesian / Libraries & Embedding

NLP in Urdu / Datasets

NLP in Urdu / Libraries

NLP in Persian / Libraries

  • Hazm

    Persian NLP Toolkit

  • Parsivar

    : A Language Processing Toolkit for Persian

  • Perke

    : Perke is a Python keyphrase extraction package for Persian language. It provides an end-to-end keyphrase extraction pipeline in which each component can be easily modified or extended to develop new models

  • Perstem

    : Persian stemmer, morphological analyzer, transliterator, and partial part-of-speech tagger

  • ParsiAnalyzer

    : Persian Analyzer For Elasticsearch

  • virastar

    : Cleaning up Persian text!

NLP in Persian / Datasets

  • Bijankhan Corpus

    : Bijankhan corpus is a tagged corpus that is suitable for natural language processing research on the Persian (Farsi) language. This collection is gathered form daily news and common texts. In this collection all documents are categorized into different subjects such as political, cultural and so on. Totally, there are 4300 different subjects. The Bijankhan collection contains about 2.6 millions manually tagged words with a tag set that contains 40 Persian POS tags

  • Uppsala Persian Corpus (UPC)

    : Uppsala Persian Corpus (UPC) is a large, freely available Persian corpus. The corpus is a modified version of the Bijankhan corpus with additional sentence segmentation and consistent tokenization containing 2,704,028 tokens and annotated with 31 part-of-speech tags. The part-of-speech tags are listed with explanations in

  • Large-Scale Colloquial Persian

    : Large Scale Colloquial Persian Dataset (LSCP) is hierarchically organized in asemantic taxonomy that focuses on multi-task informal Persian language understanding as a comprehensive problem. LSCP includes 120M sentences from 27M casual Persian tweets with its dependency relations in syntactic annotation, Part-of-speech tags, sentiment polarity and automatic translation of original Persian sentences in English (EN), German (DE), Czech (CS), Italian (IT) and Hindi (HI) spoken languages. Learn more about this project at

  • ArmanPersoNERCorpus

    : The dataset includes 250,015 tokens and 7,682 Persian sentences in total. It is available in 3 folds to be used in turn as training and test sets. Each file contains one token, along with its manually annotated named-entity tag, per line. Each sentence is separated with a newline. The NER tags are in IOB format

  • FarsiYar PersianNER

    : The dataset includes about 25,000,000 tokens and about 1,000,000 Persian sentences in total based on . The NER tags are in IOB format. More than 1000 volunteers contributed tag improvements to this dataset via web panel or android app. They release updated tags every two weeks

  • PERLEX

    : The first Persian dataset for relation extraction, which is an expert translated version of the “Semeval-2010-Task-8” dataset. Link to the relevant publication

  • Persian Syntactic Dependency Treebank

    : This treebank is supplied for free noncommercial use. For commercial uses feel free to contact us. The number of annotated sentences is 29,982 sentences including samples from almost all verbs of the Persian valency lexicon

  • Uppsala Persian Dependency Treebank (UPDT)

    : Dependency-based syntactically annotated corpus

  • Hamshahri

    : Hamshahri collection is a standard reliable Persian text collection that was used at Cross Language Evaluation Forum (CLEF) during years 2008 and 2009 for evaluation of Persian information retrieval systems

NLP in Ukrainian

  • awesome-ukrainian-nlp

    a curated list of Ukrainian NLP datasets, models, etc

  • UkrainianLT

    another curated list with a focus on machine translation and speech processing

NLP in Hungarian

  • awesome-hungarian-nlp

    : A curated list of free resources dedicated to Hungarian Natural Language Processing

NLP in Portuguese

  • Portuguese-nlp

    a List of resources and tools developed with focus on Portuguese

Other Languages

  • pymorphy2

    Russian: - a good pos-tagger for Russian

  • ICU Tokenizer

    Asian Languages: Thai, Lao, Chinese, Japanese, and Korean implementation in ElasticSearch

  • CLTK

    Ancient Languages: : The Classical Language Toolkit is a Python library and collection of texts for doing NLP in ancient languages

  • NLPH_Resources

    Hebrew: - A collection of papers, corpora and linguistic resources for NLP in Hebrew

More related projects

Add a GitHub project

Missing a project or an awesome list? Paste its GitHub URL and we fetch it right away.