Awesome Lists

awesome-hungarian-nlp

by oroszgy

awesome listpushed almost 3 years ago

A curated list of NLP resources for Hungarian

AI summary

NLP toolkit

A curated collection of NLP resources and tools for Hungarian language processing.

stars
227
forks
18
watching
20
awesome lists
2
entries
159
View on GitHuboroszgy.gitbook.io/awesome-hungarian-nlp-resources

Embed the badge

Show how many awesome lists link to your project. The count updates automatically.

Awesome Lists badge
Markdown
[![Awesome Lists Badge](https://awesome.facts.dev/shield/oroszgy/awesome-hungarian-nlp/links.svg)](https://awesome.facts.dev/awesome/oroszgy/awesome-hungarian-nlp)
HTML
<a href="https://awesome.facts.dev/awesome/oroszgy/awesome-hungarian-nlp"><img src="https://awesome.facts.dev/shield/oroszgy/awesome-hungarian-nlp/links.svg" alt="Awesome Lists Badge" /></a>
Image URL
https://awesome.facts.dev/shield/oroszgy/awesome-hungarian-nlp/links.svg

What's in the list

159 links in 25 sections, with live GitHub stats.activeno commit in 2y

Tools / Word tokenization, sentence splitting

  • huntoken

    πŸ‘ŒπŸš€πŸ’― Hungarian word and sentence splitter

  • quntoken

    πŸ‘ŒπŸš€πŸ’― New Hungarian tokenizer based on quex, huntoken

Tools / Morphology

  • emMorph (Humor)

    πŸ’― Hungarian morphological analyzer based on Humor

  • emMorphPy

    πŸ‘ŒπŸ’―A wrapper, a lemmatizer and REST API implemented in Python for emMorph (Humor) Hungarian morphological analyzer

  • hunmorph

    πŸš€πŸ’― is an open source tool and programming library for spell-checking, stemming and morphological analysing of agglutinative, german and other languages

  • hunmorph-foma

    πŸš€πŸ’― Hungarian morpholical analyzer and generator based on hunmorph

  • hunspell

    πŸ‘ŒπŸš€πŸ’― is an open-source spell-checker, stemmer and morphological analyzer

  • lara-hungarian-nlp

    πŸ‘ŒπŸš€πŸ’― LARA is a lightweight Python NLP library for ChatBots in Hungarian

  • Lemmagen

    πŸ‘ŒπŸš€πŸ’― project aims at providing standardized open source multilingual platform for lemmatisation. ( | )

  • Simplemma

    πŸ‘ŒπŸš€πŸ’― is a simple multilingual lemmatizer for Python

Tools / PoS / Morphological taggers

  • hunpos

    πŸ‘ŒπŸš€πŸ’― Hunpos is an open source reimplementation of TnT, the well known part-of-speech tagger by Thorsten Brants

  • PurePos

    πŸ‘ŒπŸš€ Open source morphological tagger based on HunPos

  • purepos.py

    πŸ‘ŒπŸš€ Python wrapper for PurePos

Tools / Taggers / Chunkers

  • HunTag

    πŸ‘ŒπŸš€ A sequential tagger for NLP using Maximum Entropy Learning and Hidden Markov Models

  • HunTag3

    πŸ‘ŒπŸš€ Improved version of the original HunTag

  • SzegedNER

    πŸ‘ŒπŸš€πŸ’― Named Entity Recognition tool for Hungarian and English

  • DBpedia Spotlight

    πŸ‘ŒπŸš€πŸ’― DBpedia Spotlight is a tool for automatically annotating mentions of DBpedia resources in text

  • emBERT

    πŸ‘ŒπŸš€πŸ’― is an emtsv module for pre-trained Transfomer-based models. It provides tagging models based on Huggingface's transformers package

Tools / Pipelines with Hungarian NLP components

  • magyarlanc

    πŸ‘ŒπŸ’― A toolkit for the basic linguistic processing of Hungarian

  • magyarlanc_spark

    πŸ‘ŒπŸ’― Spark wrapper for magyarlanc

  • eszterland

    πŸ‘ŒπŸ’― Clojurized access to magyarlanc

  • HuSpaCy

    πŸ‘ŒπŸš€πŸ’― Industrial-strength Hungarian Natural Language Processing

  • huNLP

    πŸ‘ŒπŸ’― An experimental unified Java and REST API for magyarlanc and szegedNER

  • hunlp-GATE

    πŸ’― GATE plugin containing Hungarian NLP tools as GATE processing resources

  • Trendminer Hungarian Processing Pipeline

    πŸš€ Hungarian NLP pipeline for social media text analysis (TrendMiner project)

  • Google Syntaxnet

    πŸš€πŸ’― Neural Models of Syntax

  • UDPipe

    πŸ‘ŒπŸš€πŸ’― is a trainable pipeline for tokenization, tagging, lemmatization and dependency parsing of CoNLL-U files

  • polyglot

    πŸ‘ŒπŸš€πŸ’― is a natural language pipeline that supports massive multilingual applications

  • emtsv

    πŸ‘ŒπŸ’― is a text processing system with inter-module communication via tsv + REST API

  • Stanza

    πŸ‘ŒπŸš€πŸ’― is a Python NLP Library for Many Human Languages

  • spaCy StanfordNLP

    πŸ‘ŒπŸš€πŸ’― wraps the StanfordNLP library, so you can use Stanford's models as a spaCy pipeline

  • trankit

    πŸ‘ŒπŸš€πŸ’― A Light-Weight Transformer-based Python Toolkit for Multilingual Natural Language Processing

Tools / Syntactic parsers

  • hunpars

    πŸš€πŸ’― A rule based Hungarian syntactical analyzer

  • HunParse

    πŸš€πŸ’― An NLTK-based parser using KR-style morphological annotation

  • Anagramma Parser

    A parser based on psycholinguistics principles

  • benepar

    πŸ‘ŒπŸš€πŸ’― A high-accuracy parser with models for 11 languages, implemented in Python. Based on Constituency Parsing with a Self-Attentive Encoder from ACL 2018

Tools / Semantic analysis

  • SentimentAnalysisHUN

    πŸ‘ŒπŸš€πŸ’― is an open-source sentiment analysis tool for Hungarian language, written in Python

  • hun-date-parser

    πŸ‘ŒπŸš€πŸ’― A tool for extracting datetime intervals from Hungarian sentences and turning datetime objects into Hungarian text

  • mT5-small-HunSum-1

    SZTAKI HunSum-1 models πŸ‘ŒπŸš€πŸ’― , , ,

Tools / Other

  • emLam

    πŸ‘ŒπŸš€πŸ’― Preprocessing scripts for Hungarian Language Modeling

  • pywnxml

    πŸ‘ŒπŸš€πŸ’― Python3 API for WordNet XML (Hungarian WordNet / BalkaNet / VisDic format)

  • Hun-appointment-chatbot

    πŸ‘ŒπŸš€πŸ’― A simple Hungarian chatbot for booking an appointment using the Rasa framework

  • neural-punctuator

    πŸ‘ŒπŸš€πŸ’― Automatic punctuation restoration with BERT models for English and Hungarian

  • hunaccent

    πŸ‘ŒπŸš€πŸ’― Small Footprint Diacritic Restoration for Hungarian

  • Diacritics_restoration

    πŸš€πŸ’― Lightweight Diacritics Restoration with Dilated Convolutional Neural Networks

  • NYTK MT

    πŸ‘ŒπŸš€πŸ’― NYTK Machine translation models

  • syntax-augmentation-nmt

    πŸš€πŸ’― Syntax-based data augmentation for Hungarian-English machine translation

  • anonymizer_hu

    πŸš€πŸ’― The Hungarian anonymization tool for CURLICAT

Language models / Word embeddings

Language models / Transformer models

Datasets / Corpora

  • Hungarian Webcorpus

    With over 1.48 billion words unfiltered (589 million words fully filtered), this is by far the largest Hungarian language corpus, and unlike the Hungarian National Corpus (125 million words), it is available in its entirety under a permissive Open Content license

  • Hungarian Webcorpus 2.0

    The new version of the Hungarian Webcorpus was built from Common Crawl and includes a little over 9 billion words

  • OSCAR

    is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture. (2339 million unique words)

  • emLam

    A Language Modeling Benchmark Corpus for Hungarian, similar to the One Billion Word corpus (Chelba, 2014) for English

  • Leipzig corpora

    contains randomly selected sentences in the language of the corpus and are available in sizes from 10,000 sentences up to 1 million sentences. The sources are either newspaper texts or texts randomly collected from the web

  • web2corpus

    Automatically created multilingual web corpus

  • CC-100

    Monolingual Datasets from Web Crawl Data

  • CoNLL 2017: Automatically Annotated Raw Texts and Word Embeddings

    Automatic segmentation, tokenization and morphological and syntactic annotations of raw texts in 45 languages, generated by , together with word embeddings of dimension 100 computed from lowercased texts by

  • OpinHuBank

    OpinHuBank is a human-annotated corpus to aid the research of opinion mining and sentiment analysis in Hungarian

  • HunEmPoli

    corpus was built using pre-agenda speeches of the Hungarian National Assembly (2014-2018) and consists 764008 tokens/36475 sentences. Aspect level emotion annotation, with 39840 identified emotions, in addition, marked the keywords that evoked the emotion

  • The Hungarian forum corpus for Opinion Mining

    This database is the first one dedicated to Opinion Mining in Hungarian. The data for further processing were gathered from the posts of the forum topic of the Hungarian government portal dealing with the referendum about dual citizenship

  • Hungarian sentiment corpus (HuSent)

    is a deeply annotated Hungarian sentiment corpus. It is composed of Hungarian opinion texts written about different types of products, published on the homepage [

  • Szeged Treebank

    The Szeged Treebank is the largest fully manually annotated treebank of the Hungarian language

  • Szeged Dependency Treebank

    The Szeged Dependency Treebank is a dependency-tree format version of the Szeged Treebank

  • Hungarian Named Entity Corpora

    The Named Entity Corpus for Hungarian is a subcorpus of the Szeged Treebank, which contains full syntactic annotations done manually by linguist experts

  • KorKor Pilotcorpus

    is a gold standard corpus consisting of multiple layers such as dependency parse and coreference annotations

  • NerKor

    is a gold standard named entity annotated corpus containing 1 million tokens

  • NerKor 1.41e

    A 1M+-token Hungarian named entity dataset with ~30 entity types derived from NYTK-NerKor

  • hunNERwiki

    a silver standard corpus for Hungarian Named Entity Recognition

  • Mazsola database

    contains 28M sentences from the MNSZ1 corpus annotated with shallow syntactic analysis

  • PrevCons

    is a database of 21K hapaxes of verbs with verbal prefixes

  • Hungarian word sense disambiguated corpus

    containing 39 suitable word form samples for the purpose of word sense disambiguation

  • HunLearner

    is a learners' corpus of Hungarian containing written data from 35 students majoring in Hungarian studies at the University of Zagreb, Croatia. Texts were morphologically and syntactically analyzed by the magyarlanc tool

  • HuLU

    Hungarian Language Understanding Benchmark Kit

Datasets / Corpora / HuLU

  • HuCOLA

    Hungarian Corpus of Linguistic Acceptability

  • HuCoPA

    Hungarian Choice of Plausible Alternatives Corpus

  • HuSST

    Hungarian version of the Sentiment Treebank

  • HuWNLI

    Anaphora resolution datasets for Hungarian as an inference task

  • HuWS

    is the Hungarian set of the Winograd schemas

Datasets / Corpora

  • HuRC

    Hungarian Corpus for Reading Comprehension with Commonsense Reasoning

  • ELTE Poetry Corpus

    is a database of complete poems of 50 Hungarian canonical poets together with the sound devices of the poems and the grammatical features of words in XML format

  • ELTE Novel Corpus

    is a database of 400 Hungarian novels (with the annotation of structural units and the grammatical features of words in TEI XML format)

  • ELTE Drama Corpus

    is a database of 58 dramas (with the annotation of structural units and the grammatical features of words in TEI XML format)

  • HumSum-1

    is a dataset containing over 1.1M unique news articles with lead and other metadata

  • HAPP

    is the Hungarian translation of the

  • Hunglish Corpus

    The Hunglish Corpus is a free sentence-aligned Hungarian-English parallel corpus of about 120 million words in 4 million sentence pairs

  • SzegedParallel

    The English-Hungarian parallel corpus contains texts selected on the basis of grammatical and translational criteria

  • HunOr

    A Hungarian-Russian Parallel corpus comprises approximately 800 thousand words

  • CoNLL 2017 Shared Task Hungarian data

    Automatic segmentation, tokenization and morphological and syntactic annotations of raw texts from the Common Crawl

  • CSS10

    A Collection of Single Speaker Speech Datasets for 10 Languages including Hungarian

  • TED talks transcripts parallel corpus

    sentence aligned TED talks including Hungarian

  • TaPaCo Corpus

    is a paraphrase corpus for 73 languages, including Hungarian, extracted from the Tatoeba database

  • Duolingo STAPLE

    is a dataset of comprehensive accepted translations from English to 5 different languages, including Hungarian

  • PPDB

    is an automatically extracted database containing millions of paraphrases in 16 different languages, including Hungarian

  • OpenSubtitles Corpus

    contains movie subtitles and alignments for 62 languages, including Hungarian

  • https://opus.nlpl.eu]

    [OPUS Corpus][ is a growing collection of translated texts from the web

  • MASSIVE dataset

    is a parallel dataset of > 1M utterances across 51 languages with annotations for the Natural Language Understanding tasks of intent prediction and slot annotation

  • PWS

    is a parallel collection of the Winograd schemas in seven languages (including Hungarian)

Datasets / Linguistic resources

  • morphdb.hu

    is an open source morphological database of Hungarian, consisting of a lexicon and morphological grammar that are based on well-founded theoretical decisions

  • huwn

    Hungarian Wordnet

  • Hungarian Sentiment Lexicon

    The dictionaries were manually created on the basis of Wordnet-Affect lexicons

  • poltextLAB's sentiment lexicons

    Highly accurate sentiment lexicons for analysing news data

  • 4lang

    Concept dictionary using Eilenberg machines

  • Mazsola ISZ

    lists 500K verb frames extracted from the Mazsola database

  • Manocska

    merges verb frames existing databases

  • PrevLex

    List of phrasel verbs

  • panmorph

    Tagsets and description of Hungarian morphological analysers

  • hun_ner_checklist

    CHECKLIST diagnostic test cases for Hungarian Named Entity Recognition

Datasets / Linked Open Data

Datasets / Geo data

Academy / Journals

Academy / Conferences

  • MSZNY

    Conference on Hungarian Computational Linguistics (since 2003)

Academy / Institutes

Learning resources / Books

Learning resources / Courses

Learning resources / Tutorials

Communities

More related projects

Add a GitHub project

Missing a project or an awesome list? Paste its GitHub URL and we fetch it right away.