spanish-corpora

Spanish Corpus

A collection of unannotated Spanish text data, compiled from various sources and processed for natural language processing tasks.

Unannotated Spanish 3 Billion Words Corpora

GitHub

92 stars
4 watching
10 forks
Language: Python
last commit: almost 4 years ago
Linked from 1 awesome list

corporalinguisticsnatural-language-processingnlpspanishspanish-language

Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
crscardellino/sbwceA collection of linguistic resources and trained word embeddings for the Spanish language.45
bertez/corporaA collection of Galician language data in JSON format.2
cesine/corporaforfieldlinguisticsA collection of small datasets from various languages to test and evaluate NLP scripts3
dccuchile/spanish-word-embeddingsA collection of precomputed word embeddings for the Spanish language, derived from different corpora and computational methods.354
botcenter/spanishwordembeddingsThis project generates Spanish word embeddings using fastText on large corpora.9
universaldependencies/ud_galician-ctgThis is a collection of annotated text data for the Galician language.1
dccuchile/betoA pre-trained NLP model trained on Spanish text data using the BERT architecture490
christos-c/bible-corpusA multilingual parallel corpus created from translations of the Bible.177
cidles/pyannotationA Python library to access and manipulate linguistically annotated corpus files in various formats.16
nytud/hucopaA dataset and annotation scheme for Hungarian causal reasoning tasks.1
botcenter/spanish-sent2vecThis project trains a machine learning model to generate sentence embeddings from Spanish text data using the sent2vec algorithm.4
several27/fakenewscorpusA large dataset of news articles with labeled categories to train fake news recognition algorithms385
dav009/latinamericantextresourcesA collection of linguistic and text resources for Latin America6
poltextlab/hunempoli_corpusA manually annotated corpus for training and testing machine learning models of Aspect Based Sentiment Analysis (ABSA) in Hungarian language.0
vadno/korkor_pilotA large annotated corpus of Hungarian text with various linguistic annotations, split into development and test datasets for natural language processing tasks.2