aravec

Arabic embeddings

Provides pre-trained word embedding models for Arabic text analysis

AraVec is a pre-trained distributed word representation (word embedding) open source project which aims to provide the Arabic NLP research community with free to use and powerful word embedding models.

GitHub

395 stars
32 watching
79 forks
Language: Jupyter Notebook
last commit: over 5 years ago
arabicembedded-modelsgensimnlptext-miningword2vec

Related projects:

RepositoryDescriptionStars
dfki-interactive-machine-learning/arasifProvides sentence embeddings for Arabic languages using pre-trained word embeddings and Smooth Inverse Frequency algorithm5
alexandres/lexvecAn implementation of a word embedding model that uses character n-grams and achieves state-of-the-art results in multiple NLP tasks803
wikipedia2vec/wikipedia2vecA tool for learning vector representations of words and entities from Wikipedia text data.946
galuhsahid/indonesian-word-embeddingDemonstrates word embedding in Indonesian language using pre-trained Word2vec models20
tca19/dict2vecA framework to learn word embeddings using lexical dictionaries115
satwikkottur/visualword2vecLearning word embeddings from abstract images to improve language understanding19
vefstathiou/so_word2vecThis is a word embedding model trained on Stack Overflow posts for use in natural language processing tasks.40
botcenter/spanish-sent2vecThis project trains a machine learning model to generate sentence embeddings from Spanish text data using the sent2vec algorithm.4
auspicious3000/contentvecAn implementation of a self-supervised speech representation model using PyTorch and disentangled speaker embeddings471
alexrutherford/arabic_nlpTools for normalizing and deriving sentiment from Arabic text26
hassygo/charngram2vecA repository providing a re-implementation of character n-gram embeddings for pre-training in natural language processing tasks23
hit-scir/elmoformanylangsProvides pre-trained ELMo representations for multiple languages to improve NLP tasks.1,462
cod3licious/conecA library for training and evaluating a type of word embedding model that extends the original Word2Vec algorithm20
artetxem/vecmapAn implementation of cross-lingual word embedding mappings using unsupervised learning methods648
botcenter/spanishwordembeddingsThis project generates Spanish word embeddings using fastText on large corpora.9