awesome-hungarian-nlp
by oroszgy
A curated list of NLP resources for Hungarian
AI summary
NLP toolkit
A curated collection of NLP resources and tools for Hungarian language processing.
- stars
- 227
- forks
- 18
- watching
- 20
- awesome lists
- 2
- entries
- 159
- #awesome
- #awesome-list
- #computational-linguistics
- #corpus
- #corpus-linguistics
- #dataset
- #hungarian
- #hungarian-language
- #information-extraction
- #information-retrieval
- #named-entity-recognition
- #natural-language-processing
- #natural-language-understanding
- #nlp
- #nlp-resources
- #nlu
- #opinion-mining
- #parser
- #tagger
- #text-mining
What's in the list
159 links in 25 sections, with live GitHub stats.activeno commit in 2y
Tools / Word tokenization, sentence splitting
Tools / Morphology
emMorph (Humor)
π― Hungarian morphological analyzer based on Humor
emMorphPy
ππ―A wrapper, a lemmatizer and REST API implemented in Python for emMorph (Humor) Hungarian morphological analyzer
- hunmorph
ππ― is an open source tool and programming library for spell-checking, stemming and morphological analysing of agglutinative, german and other languages
hunmorph-foma
ππ― Hungarian morpholical analyzer and generator based on hunmorph
- hunspell
πππ― is an open-source spell-checker, stemmer and morphological analyzer
lara-hungarian-nlp
πππ― LARA is a lightweight Python NLP library for ChatBots in Hungarian
- Lemmagen
πππ― project aims at providing standardized open source multilingual platform for lemmatisation. ( | )
Simplemma
πππ― is a simple multilingual lemmatizer for Python
Tools / PoS / Morphological taggers
- hunpos
πππ― Hunpos is an open source reimplementation of TnT, the well known part-of-speech tagger by Thorsten Brants
PurePos
ππ Open source morphological tagger based on HunPos
purepos.py
ππ Python wrapper for PurePos
Tools / Taggers / Chunkers
HunTag
ππ A sequential tagger for NLP using Maximum Entropy Learning and Hidden Markov Models
HunTag3
ππ Improved version of the original HunTag
- SzegedNER
πππ― Named Entity Recognition tool for Hungarian and English
DBpedia Spotlight
πππ― DBpedia Spotlight is a tool for automatically annotating mentions of DBpedia resources in text
emBERT
πππ― is an emtsv module for pre-trained Transfomer-based models. It provides tagging models based on Huggingface's transformers package
Tools / Pipelines with Hungarian NLP components
- magyarlanc
ππ― A toolkit for the basic linguistic processing of Hungarian
magyarlanc_spark
ππ― Spark wrapper for magyarlanc
eszterland
ππ― Clojurized access to magyarlanc
HuSpaCy
πππ― Industrial-strength Hungarian Natural Language Processing
huNLP
ππ― An experimental unified Java and REST API for magyarlanc and szegedNER
hunlp-GATE
π― GATE plugin containing Hungarian NLP tools as GATE processing resources
Trendminer Hungarian Processing Pipeline
π Hungarian NLP pipeline for social media text analysis (TrendMiner project)
- Google Syntaxnet
ππ― Neural Models of Syntax
- UDPipe
πππ― is a trainable pipeline for tokenization, tagging, lemmatization and dependency parsing of CoNLL-U files
- polyglot
πππ― is a natural language pipeline that supports massive multilingual applications
emtsv
ππ― is a text processing system with inter-module communication via tsv + REST API
Stanza
πππ― is a Python NLP Library for Many Human Languages
spaCy StanfordNLP
πππ― wraps the StanfordNLP library, so you can use Stanford's models as a spaCy pipeline
trankit
πππ― A Light-Weight Transformer-based Python Toolkit for Multilingual Natural Language Processing
Tools / Syntactic parsers
- hunpars
ππ― A rule based Hungarian syntactical analyzer
HunParse
ππ― An NLTK-based parser using KR-style morphological annotation
Anagramma Parser
A parser based on psycholinguistics principles
benepar
πππ― A high-accuracy parser with models for 11 languages, implemented in Python. Based on Constituency Parsing with a Self-Attentive Encoder from ACL 2018
Tools / Semantic analysis
SentimentAnalysisHUN
πππ― is an open-source sentiment analysis tool for Hungarian language, written in Python
hun-date-parser
πππ― A tool for extracting datetime intervals from Hungarian sentences and turning datetime objects into Hungarian text
- mT5-small-HunSum-1
SZTAKI HunSum-1 models πππ― , , ,
Tools / Other
emLam
πππ― Preprocessing scripts for Hungarian Language Modeling
pywnxml
πππ― Python3 API for WordNet XML (Hungarian WordNet / BalkaNet / VisDic format)
Hun-appointment-chatbot
πππ― A simple Hungarian chatbot for booking an appointment using the Rasa framework
neural-punctuator
πππ― Automatic punctuation restoration with BERT models for English and Hungarian
hunaccent
πππ― Small Footprint Diacritic Restoration for Hungarian
Diacritics_restoration
ππ― Lightweight Diacritics Restoration with Dilated Convolutional Neural Networks
NYTK MT
πππ― NYTK Machine translation models
syntax-augmentation-nmt
ππ― Syntax-based data augmentation for Hungarian-English machine translation
anonymizer_hu
ππ― The Hungarian anonymization tool for CURLICAT
Language models / Word embeddings
- FasText Wikipedia
pre-trained word vectors for 90 languages, trained on Wikipedia using fastText
- FasText Common Crawl & Wikipedia
pre-trained word vectors for 157 languages, trained on Wikipedia and the Common Crawl using fastText's CBOW model
FastText_multilingual
Multilingual word vectors in 78 languages
- polyglot vectors
polgyglot embeddings on Wikipedia
wordvectors
Pre-trained word2vec and fasttext word vectors on wikipedia of 30+ languages
- hunembed0.0
A word2vec word embedding trained on the concatenation of the Hungarian Webcorpus and the Hungarian National Corpus in 600 dimensions with a cut-off of 10 words
- Szeged word vectors
Word embeddings (word2vec & fasttext) for Hungarian trained on 4.3 billion tokens
- questions-words-hu
Hungarian analogical questions following Mikolov et al
Conceptnet Numberbatch
Conceptnet numbermatch multi- and cross-lingual semantic word embeddings
- BytePair Embeddings
pretrained Subword Embeddings, downloadable in many formats
- HuSpaCy 300d
300d Floret embeddings trained on the Hungarian Webcorpus 2.0
- HuSpaCy 100d
100d Floret embeddings trained on the Hungarian Webcorpus 2.0
ELMo Representations
Deep contextualized word representation trained for many languages
Language models / Transformer models
- huBERT
Hungarian BERT base models trained on Webcorpus 2.0 and the Hungarian Wikipedia
- HIL* Transformer models
Pretrained transformer models provided by HILANCO
- PULI-BERT-Large
is a Hungarian BERT large model based on MegatronBERT
- PULI-GPT-2
is a Hungarian GPT-2 model
- PULI-GPT-3SX
is a Hungarian GPT-NeoX model (6.7 billion parameter)
Datasets / Corpora
- Hungarian Webcorpus
With over 1.48 billion words unfiltered (589 million words fully filtered), this is by far the largest Hungarian language corpus, and unlike the Hungarian National Corpus (125 million words), it is available in its entirety under a permissive Open Content license
- Hungarian Webcorpus 2.0
The new version of the Hungarian Webcorpus was built from Common Crawl and includes a little over 9 billion words
- OSCAR
is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture. (2339 million unique words)
- emLam
A Language Modeling Benchmark Corpus for Hungarian, similar to the One Billion Word corpus (Chelba, 2014) for English
- Leipzig corpora
contains randomly selected sentences in the language of the corpus and are available in sizes from 10,000 sentences up to 1 million sentences. The sources are either newspaper texts or texts randomly collected from the web
- web2corpus
Automatically created multilingual web corpus
- CC-100
Monolingual Datasets from Web Crawl Data
- CoNLL 2017: Automatically Annotated Raw Texts and Word Embeddings
Automatic segmentation, tokenization and morphological and syntactic annotations of raw texts in 45 languages, generated by , together with word embeddings of dimension 100 computed from lowercased texts by
- OpinHuBank
OpinHuBank is a human-annotated corpus to aid the research of opinion mining and sentiment analysis in Hungarian
HunEmPoli
corpus was built using pre-agenda speeches of the Hungarian National Assembly (2014-2018) and consists 764008 tokens/36475 sentences. Aspect level emotion annotation, with 39840 identified emotions, in addition, marked the keywords that evoked the emotion
- The Hungarian forum corpus for Opinion Mining
This database is the first one dedicated to Opinion Mining in Hungarian. The data for further processing were gathered from the posts of the forum topic of the Hungarian government portal dealing with the referendum about dual citizenship
- Hungarian sentiment corpus (HuSent)
is a deeply annotated Hungarian sentiment corpus. It is composed of Hungarian opinion texts written about different types of products, published on the homepage [
- Szeged Treebank
The Szeged Treebank is the largest fully manually annotated treebank of the Hungarian language
- Szeged Dependency Treebank
The Szeged Dependency Treebank is a dependency-tree format version of the Szeged Treebank
- Hungarian Named Entity Corpora
The Named Entity Corpus for Hungarian is a subcorpus of the Szeged Treebank, which contains full syntactic annotations done manually by linguist experts
KorKor Pilotcorpus
is a gold standard corpus consisting of multiple layers such as dependency parse and coreference annotations
NerKor
is a gold standard named entity annotated corpus containing 1 million tokens
NerKor 1.41e
A 1M+-token Hungarian named entity dataset with ~30 entity types derived from NYTK-NerKor
- hunNERwiki
a silver standard corpus for Hungarian Named Entity Recognition
- Mazsola database
contains 28M sentences from the MNSZ1 corpus annotated with shallow syntactic analysis
PrevCons
is a database of 21K hapaxes of verbs with verbal prefixes
- Hungarian word sense disambiguated corpus
containing 39 suitable word form samples for the purpose of word sense disambiguation
- HunLearner
is a learners' corpus of Hungarian containing written data from 35 students majoring in Hungarian studies at the University of Zagreb, Croatia. Texts were morphologically and syntactically analyzed by the magyarlanc tool
HuLU
Hungarian Language Understanding Benchmark Kit
Datasets / Corpora / HuLU
Datasets / Corpora
- HuRC
Hungarian Corpus for Reading Comprehension with Commonsense Reasoning
ELTE Poetry Corpus
is a database of complete poems of 50 Hungarian canonical poets together with the sound devices of the poems and the grammatical features of words in XML format
ELTE Novel Corpus
is a database of 400 Hungarian novels (with the annotation of structural units and the grammatical features of words in TEI XML format)
ELTE Drama Corpus
is a database of 58 dramas (with the annotation of structural units and the grammatical features of words in TEI XML format)
- HumSum-1
is a dataset containing over 1.1M unique news articles with lead and other metadata
HAPP
is the Hungarian translation of the
- Hunglish Corpus
The Hunglish Corpus is a free sentence-aligned Hungarian-English parallel corpus of about 120 million words in 4 million sentence pairs
- SzegedParallel
The English-Hungarian parallel corpus contains texts selected on the basis of grammatical and translational criteria
- HunOr
A Hungarian-Russian Parallel corpus comprises approximately 800 thousand words
- CoNLL 2017 Shared Task Hungarian data
Automatic segmentation, tokenization and morphological and syntactic annotations of raw texts from the Common Crawl
CSS10
A Collection of Single Speaker Speech Datasets for 10 Languages including Hungarian
- TED talks transcripts parallel corpus
sentence aligned TED talks including Hungarian
- TaPaCo Corpus
is a paraphrase corpus for 73 languages, including Hungarian, extracted from the Tatoeba database
- Duolingo STAPLE
is a dataset of comprehensive accepted translations from English to 5 different languages, including Hungarian
- PPDB
is an automatically extracted database containing millions of paraphrases in 16 different languages, including Hungarian
- OpenSubtitles Corpus
contains movie subtitles and alignments for 62 languages, including Hungarian
- https://opus.nlpl.eu]
[OPUS Corpus][ is a growing collection of translated texts from the web
MASSIVE dataset
is a parallel dataset of > 1M utterances across 51 languages with annotations for the Natural Language Understanding tasks of intent prediction and slot annotation
PWS
is a parallel collection of the Winograd schemas in seven languages (including Hungarian)
Datasets / Linguistic resources
- morphdb.hu
is an open source morphological database of Hungarian, consisting of a lexicon and morphological grammar that are based on well-founded theoretical decisions
huwn
Hungarian Wordnet
- Hungarian Sentiment Lexicon
The dictionaries were manually created on the basis of Wordnet-Affect lexicons
poltextLAB's sentiment lexicons
Highly accurate sentiment lexicons for analysing news data
4lang
Concept dictionary using Eilenberg machines
- Mazsola ISZ
lists 500K verb frames extracted from the Mazsola database
Manocska
merges verb frames existing databases
PrevLex
List of phrasel verbs
panmorph
Tagsets and description of Hungarian morphological analysers
hun_ner_checklist
CHECKLIST diagnostic test cases for Hungarian Named Entity Recognition
Datasets / Linked Open Data
huwn.rdf
Hungarian WordNet in RDF format for the Linked Open Data cloud
- Conceptnet
An open, multilingual knowledge graph (with partial Hungarian support)
Datasets / Geo data
- OpenStreetMap(OSM)
In the keys, the
Natural-earth-vector
( imported from wikidata labels)
- Who's On First
is a gazetteer of places (with )
Datasets / Speech related data
Academy / Journals
Academy / Conferences
- MSZNY
Conference on Hungarian Computational Linguistics (since 2003)
Academy / Institutes
Learning resources / Books
Learning resources / Courses
Learning resources / Tutorials
Communities
- KeresΕ vilΓ‘g
Official blog of Precognox Inc
Other Hungarian related resource collections
EENLP
The broad index of NLP resources for Eastern European languages
Nothing in this list matches your filter.
Featured in 2 awesome lists
Each link jumps to the spot where the list mentions awesome-hungarian-nlp.
More related projects
gianlucabertani/machinelearning37
facebookresearch/fasttext26K
facebookresearch/muse3.2K
dccuchile/spanish-word-embeddings354
pymorphy2/pymorphy21.1K
amakukha/stemmers_ukrainian28
curiosity-ai/catalyst752
web64/norwegian-nlp-resources178
helsinki-nlp/ukrainianlt30
pawangeek/deep-nlp-resources73
jdidion/biotools590
jameslavin/my_tech_resources312