HuLU

Language datasets

A collection of linguistic datasets and benchmarks for natural language understanding tasks

Hungarian Language Understanding Benchmark Kit

GitHub

8 stars
3 watching
0 forks
last commit: about 2 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
nytud/huwsA dataset of manually curated Hungarian sentences with ambiguous wordings that require world knowledge and reasoning for resolution.1
nytud/husstA dataset of annotated sentences for training and evaluating sentiment analysis models in the Hungarian language.1
nytud/hucolaA collection of 9,076 annotated sentences in Hungarian to evaluate linguistic acceptability and grammaticality1
nytud/happA dataset of Hungarian translations of human-language examples to test anaphora resolution algorithms1
nytud/huwnliA dataset and toolset for Hungarian anaphora resolution in natural language inference tasks0
nytud/pwsA collection of parallel corpora of Winograd schemata in multiple languages0
nytud/hunlp-gateA collection of Hungarian NLP tools integrated as GATE processing resources8
nytud/panmorphHarmonized tagset and annotation scheme for Hungarian morphological analysers4
nytud/machine-translationProvides machine translation models and a demo site for Hungarian language translations5
xuefuzhao/instructionwildCreating a large-scale user-based instruction dataset for natural language processing research and development455
alexa/massiveA collection of tools and modeling code for a large multilingual Natural Language Understanding dataset541
turkunlp/wikibertProvides pre-trained language models derived from Wikipedia texts for natural language processing tasks34
karthikncode/nlp-datasetsA curated list of Natural Language Processing datasets used to train and evaluate NLP models.919
nytud/hadifogoly-adatbazisAn attempt to transcribe Cyrillic text into Hungarian script for a large dataset of WWII prisoner-of-war records23
novakat/nytk-nerkor-cars-ontonotesppA large annotated dataset of Hungarian text with over 30 entity types derived from various sources and formats.1