Awesome Lists

awesome-danish

by fnielsen

awesome listpushed almost 2 years ago

A curated list of awesome resources for Danish language technology

AI summary

Danish NLP dataset

A curated collection of Danish language resources and datasets for natural language processing tasks.

stars
168
forks
18
watching
16
awesome lists
2
entries
125
View on GitHub

Embed the badge

Show how many awesome lists link to your project. The count updates automatically.

Awesome Lists badge
Markdown
[![Awesome Lists Badge](https://awesome.facts.dev/shield/fnielsen/awesome-danish/links.svg)](https://awesome.facts.dev/awesome/fnielsen/awesome-danish)
HTML
<a href="https://awesome.facts.dev/awesome/fnielsen/awesome-danish"><img src="https://awesome.facts.dev/shield/fnielsen/awesome-danish/links.svg" alt="Awesome Lists Badge" /></a>
Image URL
https://awesome.facts.dev/shield/fnielsen/awesome-danish/links.svg

What's in the list

125 links in 31 sections, with live GitHub stats.activeno commit in 2y

Data / Corpora

  • Danish Gigaword

    Collection of 10^12 words of Danish text. Described in ( )

  • Danish review dataset

    Trustpilot-crawled dataset by Alessandro Gianfelici with 44,085 reviews

  • OSCAR

    Danish corpus derived from the Common Crawl corpus. Described in ( )

Data / Corpora / CLARIN-DK-UCPH

Data / Corpora

  • DanFEVER

    Danish text corpus with over 6'400 claims and support. Described in ( )

  • DanNet

    wordnet with usage examples. The usage examples have been used for word sense disambiguation, see

  • SemDaX

    POS-tagged (only adjectives, nouns and verbs), super sense tagged and BIO-tagged sentences. For educational, teaching or research purposes only

  • NOMCO

    "an annotated multimodal collection of conversational Danish". Apparently not directly available for download. [ ]

  • Danish Propbank

    commercial resource with 87,000 tokens annotated with morphosyntactic, VerbNet classes and semantic roles

  • Danish Dependency Treebank v. 1.0

    Matthias Trautner Kromann et al.'s dependency annotation of some texts from PAROLE

  • Mr. Bean corpus

    Small Danish-Italian corpus with written and spoken retelling (of Mr Bean episodes) and argumentative text (about smoking). Possibly described in

  • Køge Corpus

    Danish-Turkish transcribed corpus by Jens Normann Jørgensen

  • Danske taler

    Collection of Danish speeches. API available at

  • DKhate

    corpus of 3600 hate speech from Twitter and Reddits as well as news comments. Described in ( )

  • Scholia

    DaNewsroom - Danish summarization dataset. Probably to appear in 2020. Described in ( )

Data / Corpora / Wikipedia

  • wiki40b/da

    Clean-up text from Danish Wikipedia. Described in . ( )

Data / Corpora

  • XED

    emotion annotated movie subtitles. Described in ( )

  • DaN+

    annotated for nested named entities on top of the entire Danish Universal Dependencies (UD_Danish-DDT) and 3 new web domains and includes lexical normalization. Described in

  • WikiANN

    Named entity annotated corpus. Described in ( )

  • Corona Dataset

    Question dataset from Certainly annotated for domain and intent

Data / Parallel corpora

  • Europarl

    parallel sentences between Danish and English from the European Parlament

  • ITU Faroese Pairs Dataset

    Faroese-Danish parallel text. Described in ( )

  • JW300

    "a parallel corpus of over 300 languages with around 100 thousand parallel sentences per language pair on average"

  • OpenSubtitles2018

    Parallel corpus from movie and tv subtitles. Described in

  • Tatoeba

    Sentences

  • WikiMatrix

    , parallel sentences from Wikipedias. 1620 language pairs, including Danish

Data / Spoken language corpora

  • CoRal

    Danish Conversational and Read-aloud Dataset

  • DanPASS

    Described in ( )

  • LANCHART

    Centre for Language Change In Real Time. Various audio recordings. Whether the data is available is not immediately apparent. Described in, e.g., ( )

  • Common Voice

    Crowdsourced multilingual annotated speech dataset. As of March 2023, 11 hours of validated speech are distributed. Sentences can be entered collaboratively at . Common Voice is described in ( )

  • FT Speech

    Described in ( )

Data / Spoken language corpora / NST

  • NST-speech-22khz

    A 22kHz speech corpus compiled by Nordisk Språkteknologi and made available by the Norwegian Library Service. The speech genre is dictation

  • NST-speech-16kHz

    A 16kHz speech corpus compiled by Nordisk Språkteknologi and made available by the Norwegian Library Service. The speech genre is read-aloud and the text is phonetically balanced. Designed for ASR training and testing

  • NST-speech-44kHz

    A 44kHz speech corpus compiled by Nordisk Språkteknologi and made available by the Norwegian Library Service. Designed for speech synthesis

Data / Spoken language corpora

  • VoxLingua107

    28 hours audio with unannotated Danish speech sampled from YouTube videos. Described in ( )

  • VoxPopuli

    Speech from the European Parliament including 13'600 hours of unannotated Danish. Described in ( )

  • Wikimedia Commons Audio files of Danish language

    Recordings of readings of articles from the Danish Wikipedia, Danish words and a few Danish literary works

Data / Dictionaries and ontologies

Data / Dictionaries and ontologies / Retskrivningsordbogen

  • Excerpt

    Lexemes, word classes and inflections. in the CSF format available. Full list presumably available upon request

Data / Dictionaries and ontologies

  • Stavekontrolden

    word list with 160,132 Danish words. Used, e.g., for spelling suggestion in LibreOffice. Licensed under GPL, LPGL, and MPL

  • The Concise Danish Dictionary

    /The Comprehensive Danish Dictionary/Den Store Danske Ordliste (DSDO), word list created by Skåne Sjælland Linux User Group and distributed under a GPL license

  • Interactive Terminology for Europe

    (IATE) - European Union terminology database. October 2020 version contains over 500,000 Danish terms

  • The Danish FrameNet Lexicon

    , 40,267 lines resource containing 5,300 verbs and 6,490 verbal nouns

  • 1,290,000 lexemes

    Wikidata lexemes - structured database with metadata about lexemes, their forms and their sense. Over including in April 2024

Data / Dictionaries and ontologies / 1,290,000 lexemes

Data / Dictionaries and ontologies

  • NST-ngrams

    A N-gram frequency list compiled by Nordisk Språkteknologi from newspaper text and made available by the Norwegian Library Service. Can be compiled to an n-gram LM with SRILM

  • AFINN

    Danish lexicons annotated for sentiment

  • concreteness-estimates-da

    Bill D. Thompson's concreteness estimates for Danish words, as detailed in ( )

  • SAM lexicon

    sentiment analysis word list extended from AFINN to 4275 lines. Described in

  • Danish Swadesh List

    List of Danish words of basic concepts from The Rosetta Project

  • Sketch Engine

    cloud service with wordlists, thesearus, collocations, n-grams etc. Free for academic use in the European Union and paid service for commercial use

Data / Word sets

  • Danish-Similarity-Dataset

    Similarity scores for 99 Danish word pairs by Nina Schneidermann and Bolette Sandford Pedersen. Also available in

  • Wordsim353-da

    Danish translation by Finn Årup Nielsen of the English Wordsim353 English word pair set. Also available in

  • Clinical similarity dataset

    289 word pairs score for similarity

  • Four words

    100 odd-one-out sets of 4 words or phrases

Data / Embeddings

  • cc.da.300

    ( ) - fastText-trained embedding on Danish part of and Danish Wikipedia. Read more about the method in ( )

  • wiki.da

    ( ) - fastText-trained embedding on Danish Wikipedia. Read more about the method in ( )

  • Byte-Pair Encoding embedding

    Gensim-based subword embedding. A large number of Danish embeddings are available. They differ in the size of the vocabulary (from 1000 to 200000) and subspace dimensions (from 25 to 300)

  • NLPL word embeddings repository

    NLPL word embeddings repository by Language Technology Group at the University of Oslo. Two Danish embedding models as of November 2020

Data / Embeddings / NLPL word embeddings repository

Data / Embeddings

Data / Neural text models

  • A-ttack

    Ælæctra-based model for detection of "textual attacks" developed by . Related to the Ha-te model

  • Danish BERT

    Certainly's (Botxo/Møllerhøj) Weights for a BERT trained on a large Danish corpora

  • Danish ELECTRA

    Philip Tamimi-Sarnikowski's Danish ELECTRA model. Available in the transformer library

  • daT5-summariser

    Danish abstractive summarisation of news articles based on mT5-base

  • ConvBERT

    Philip Tamimi-Sarnikowski's model

  • Danish ELMo on OSCAR

    (Link does not work as of December 2020)

  • Ha-te

    Hate speech detection based on Ælæctra developed by . Related to the A-ttack model

  • mfaq

    Multilingual FAQ retrieval model. Described in ( )

  • Ælæctra

    Malte Højmark-Bertelsen's Danish Gigaword-trained Electra-based model

  • Multilingual sentence transformers

    Pre-trained multilingual sentence transformers,

  • wiki40b-lm-da

    language model trained on Danish from Wiki40B dataset

  • WikiBERT

    BERT model for many languages, including Danish. Described in ( )

Data / Neural speech models

Tools / Lemmatization

  • Lemmy

    Lemmatizer for Danish in Python

  • cstlemma

    lemmatiser

  • spaCy

    Python-based package with lemmatization

Tools / Punctuation

  • punctfix

    "Adds punctuation and capitalization for a given text."

Tools / Named entity recognition

  • ScandiNER

    Scandinavian named entity recognition, achieving state-of-the-art performance in Danish, Norwegian (both Bokmål and Nynorsk), Swedish, Icelandic and Faroese

  • DaLUKE

    Danish named entity recognition based on LUKE. Described in

  • spaCy

    Python-based named entity extraction

  • daner

    Named entity extraction from ITU NLP. Described in ( )

  • flair+danlp ner-tagger

    Flair NER tagger trained by the Alexandra Institute

Tools / Entity linking

Tools / Sentiment analysis

  • afinn

    Python package with AFINN Danish lexicon annotated for sentiment, also installable with

  • Hisia

    Python package with pre-trained machine-learning based Danish sentiment analysis by Prayson Wilfred Daniel

  • senda

    Python package with transformer-based sentiment analysis from Ekstra Bladet Analyse with as of 2021 on one dataset

  • Sentida

    R package With Danish sentiment lexicon and handling of, e.g., negation. Detailed in ( )

Tools / Automatic Speech Recognition

  • danspeech

    DeepSpeech2-based Danish speech recognition in Python

  • kaldi-sprakbanken

    A recipe for training state-of-the-art(2017) speech recogniser for Danish based on the 16kHz NST database

Tools / Speech Synthesis (text-to-speech)

  • espeak

    An open-source speech synthesis program for ~56 languages including Danish. eSpeak can also be used as a grapheme-to-phoneme converter and was used to create the Danish Kaldi recipe

  • ResponsiveVoice

    Commercial Web-based (Javascript-based) text-to-speech synthesis for a number of languages, including Danish. The commercial service is currently free for limited and non-commercial use

  • Google Cloud Text-to-Speech

    Commercial Web-based text-to-speech synthesis for a number of languages, including Danish

  • Amazon Polly

    Commercial Web-based text-to-speech synthesis for a number of languages, including Danish. Part of Amazon's commercial AWS services. Female and male voices are available as examples. Limited unregistered free service available at

Tools / Fundamental processing

  • DaNLP

    "a repository for Natural Language Processing resources for the Danish Language."

  • dapipe

    Danish UD-pipe: tokenisation, lemmatisation, PoS tagging, morphology, dependencies

  • UDPipe

    Non-language specific version of dapipe. Newer version of the Danish-DDT model than that which is offered by dapipe is available at

  • DKIE

    GATE pipeline including wrapped Danish models for Stanford CoreNLP

  • StanfordNLP

    . Python software package for dependency parsing, including tokenization, lemmatization and part-of-speech tagging. A pre-trained model for Danish is available

  • bornholmsk

    Datasets and embeddings for the Bornholmsk dialect

  • spaCy

    Python-based natural language processing package

  • dacy

    Danish spaCy pipeline

Competitions

Benchmarks

  • Danoliterate

    Overview of the performance of language models on a range of individual benchmarks

  • ScandEval

    Overview of the performance of language models on a range of individual benchmark, Danish as well as other Germanic languages

Resources about resources

More related projects

Add a GitHub project

Missing a project or an awesome list? Paste its GitHub URL and we fetch it right away.