awesome-ukrainian-nlp
by osyvokon
Curated list of Ukrainian natural language processing (NLP) resources (corpora, pretrained models, libriaries, etc.)
AI summary
NLP toolkit
A curated collection of Ukrainian NLP resources for research and development.
- stars
- 168
- forks
- 14
- watching
- 11
- awesome list
- 1
- entries
- 78
What's in the list
78 links in 21 sections, with live GitHub stats.activeno commit in 2y
1. Datasets / Corpora / Monolingual
- Malyuk
— 113GB of text, compilation of UberText 2.0, OSCAR, Ukrainian News
Brown-UK
— carefully curated corpus of modern Ukrainian language with dismabiguated tokens, 1 million words
- UberText 2.0
— over 5 GB of news, Wikipedia, social, fiction, and legal texts
- OSCAR
— shuffled sentences extracted from and classified with a language detection model. Ukrainian portion of it is 28GB deduplicated
- CC-100
— documents extracted from , automatically classified and filtered. Ukrainian part is 200M sentences or 10GB of deduplicated text
mC4
— filtered CommonCrawl again, 196GB of Ukrainian text
Ukrainian Twitter corpus
Ukrainian Twitter corpus for toxic text detection
Ukrainian forums
— 250k sentences scraped from forums
- Ukrainain news headlines
— 5.2M news headlines
1. Datasets / Corpora / Parallel
- Wiki Edits
— 5M sentence edits extracted from the Ukrainian Wikipedia revision history
1. Datasets / Corpora / Labeled
- ZNO
— ~4000 questions and answers from Ukrainian External independent testing (ЗНО/ZNO)
UA-GEC
— grammatical error correction (GEC) and fluency corpus
NER-uk
— Brown-UK labeled for named entities
- Yakaboo Book Reviews
— book reviews, ratings and descriptions
Universal Dependencies
— dependency trees corpus
ua-news
— 150k news article in 5 categories
UA-SQuAD
— Ukrainian version of Stanford Question Answering Dataset
Ukrainian Winograd schema challenge (WSC) Dataset
— manually translated
Ukrainian OntoNotes Dataset
— scripts to build large silver dataset for coreference resolution
1. Datasets / Corpora / Dictionaries
ВЕСУМ
— POS tag dictionary. Can generate a list of all word forms valid for spelling
- Multilingualsentiment, includes Ukrainian
a list of positive/negative words
obscene-ukr
— profanity dictionary
Word stress dictionary
— word stress for 2.7M word forms. See
Heteronyms
— words that share the same spelling but have different meaning/pronunciation
Abbreviations
— map abbreviation to expansion
1. Datasets / Corpora / Prompts
- Aya
— crowd-sourced prompts and reference outputs. Ukrainian part is ~500 prompts
2. Tools
tree_stem
— stemmer
pymorphy2
+ — POS tagger and lemmatizer
- LanguageTool
— grammar, style and spell checker
- Stanza
— Python package for tokenization, multi-word-tokenization, lemmatization, POS, dependency parsing, NER
nlp-uk
— Tools for cleaning and normalizing texts, tokenization, lemmatization, POS, disambiguation
NLP-Cube
Python package for tokenization, sentence splitting, multi-word-tokenization, lemmatization, part-of-speech tagging and dependency parsing
3. Pretrained models / Language models
- aya-101
— massively multilingual LM, 13B parameters
- pythia-uk
— mT5 finetuned on wiki and oasst1 for chats in Ukrainian
UAlpaca
— Llama fine-tuned for instruction following on the machine-translated Alpaca dataset
XGLM
— multilingual autoregressive LM, the 4.5B checkpoint includes Ukrainian
uk4b
and - GPT-2 small, medium and large-style models trained on UberText 2.0 wikipedia, news and books
- xlm-roberta-base-uk
— truncated version of XLM-RoBERTa with only Ukrainian and English embeddings left
3. Pretrained models / Machine translation
Helsinki-NLP / OPUS-MT models
— Ukrainian to/from 25 langaguages
3. Pretrained models / Machine translation / Helsinki-NLP / OPUS-MT models
3. Pretrained models / Machine translation
M2M-100
— Ukrainian to/from 100 languages
Uk-En folktale corpus
— small sentence-aligned corpus of fairy tales
3. Pretrained models / Sequence-to-sequence models
3. Pretrained models / Named-entity recognition (NER)
3. Pretrained models / Part-of-speech tagging (POS)
3. Pretrained models / Word embeddings / fastText
- Official fastText trained on CommonCrawl and Wiki
— 157 languages, including Ukrainian
Older official fastText trained on Wiki
— 294 languages, including Ukrainian
fastText_multilingual
— 78 languages, aligned to the same vector space
- fasttext_uk (2023)
and — trained on UberText 2.0
3. Pretrained models / Word embeddings
- BPEmb: Subword Embeddings, includes Ukrainian
easy to use with
Flair
— added in 2022
3. Pretrained models / Other
- uk-punctcase
— punctuation and case restoration model based on XLM-RoBERTa-Uk
- punctuation_uk_bert
— another punctation and case restoration model based on bert-base-multilingual-cased
ukrainian-word-stress
— adds word stress
4. Paid
- LORELEI Ukrainian Representative Language Pack
Ukrainian monolingual text, Ukrainian-English parallel text, partially annotated for named entities
5. Other resources and links
Helsinki-NLP/ UkrainianLT
— another collection of links to Ukrainian language tools
egorsmkv / speech-recognition-uk
— speech recognition and text-to-speech models and datasets
6. Workshops and conferences
6. Workshops and conferences / UNLP 2023 shared task — shared task (competition) in grammatical error correction for Ukrainian
6. Workshops and conferences
Nothing in this list matches your filter.
Featured in 1 awesome list
Each link jumps to the spot where the list mentions awesome-ukrainian-nlp.