awesome-sentence-embedding

A curated list of pretrained sentence and word embedding models

Archived

GitHub

2k stars
78 watching
261 forks
Language: Python
last commit: over 5 years ago
Linked from 1 awesome list

awesomeawesome-listbertcontextualized-representationcross-lingualembedding-modelslanguage-modelnatural-languagenlppretrained-embeddingpretrained-language-modelpretrained-modelssentence-embeddingssentence-representationssubword-modelsunsupervised-learningword-embeddingswordembedding

awesome-sentence-embedding / Word Embeddings

WebVectors: A Toolkit for Building Web Interfaces for Vector Semantic Models
RusVectōrēs
Efficient Estimation of Word Representations in Vector Space
C1,525over 3 years ago
Word2Vec
Word Representations via Gaussian Embedding
Cython190over 8 years ago
A Probabilistic Model for Learning Multi-Prototype Word Embeddings
DMTK116about 10 years ago
Dependency-Based Word Embeddings
C++
word2vecf
GloVe: Global Vectors for Word Representation
C6,908almost 2 years ago
GloVe6,908almost 2 years ago
Sparse Overcomplete Word Vector Representations
C++54almost 9 years ago
From Paraphrase Database to Compositional Paraphrase Model and Back
Theano30over 10 years ago
PARAGRAM
Non-distributional Word Vector Representations
Python62about 9 years ago
WordFeat62about 9 years ago
Joint Learning of Character and Word Embeddings
C299about 6 years ago
SensEmbed: Learning Sense Embeddings for Word and Relational Similarity
SensEmbed
Topical Word Embeddings
Cython314over 8 years ago
Swivel: Improving Embeddings by Noticing What's Missing
TF77,258almost 2 years ago
Counter-fitting Word Vectors to Linguistic Constraints
Python145over 6 years ago
counter-fitting(broken)
Mixing Dirichlet Topic Models and Word Embeddings to Make lda2vec
Chainer3,152almost 5 years ago
Siamese CBOW: Optimizing Word Embeddings for Sentence Representations
Theano
Siamese CBOW
Matrix Factorization using Window Sampling and Negative Sampling for Improved Word Representations
Go803over 5 years ago
lexvec803over 5 years ago
Enriching Word Vectors with Subword Information
C++25,979over 2 years ago
fastText
Morphological Priors for Probabilistic Neural Word Embeddings
Theano52almost 10 years ago
A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks
C++23almost 3 years ago
charNgram2vec
ConceptNet 5.5: An Open Multilingual Graph of General Knowledge
Python1,296about 4 years ago
Numberbatch1,296about 4 years ago
Learning Word Meta-Embeddings
Meta-Emb(broken)
Offline bilingual word vectors, orthogonal transformations and the inverted softmax
Python1,197over 3 years ago
Multimodal Word Distributions
TF283about 7 years ago
word2gm283about 7 years ago
Poincaré Embeddings for Learning Hierarchical Representations
Pytorch1,684about 2 years ago
Context encoders as a simple but powerful extension of word2vec
Python20over 6 years ago
Semantic Specialisation of Distributional Word Vector Spaces using Monolingual and Cross-Lingual Constraints
TF64almost 9 years ago
Attract-Repel64almost 9 years ago
Learning Chinese Word Representations From Glyphs Of Characters
C30over 8 years ago
Making Sense of Word Embeddings
Python212over 5 years ago
sensegram
Hash Embeddings for Efficient Word Representations
Keras42almost 9 years ago
BPEmb: Tokenization-free Pre-trained Subword Embeddings in 275 Languages
Gensim1,189almost 2 years ago
BPEmb1,189almost 2 years ago
SPINE: SParse Interpretable Neural Embeddings
Pytorch52over 6 years ago
SPINE
AraVec: A set of Arabic Word Embedding Models for use in Arabic NLP
Gensim395over 5 years ago
AraVec395over 5 years ago
Ngram2vec: Learning Improved Word Representations from Ngram Co-occurrence Statistics
C848about 7 years ago
Dict2vec : Learning Word Embeddings using Lexical Dictionaries
C++115over 5 years ago
Dict2vec115over 5 years ago
Joint Embeddings of Chinese Words, Characters, and Fine-grained Subcharacter Components
C99about 7 years ago
Representation Tradeoffs for Hyperbolic Embeddings
Pytorch377about 3 years ago
h-MDS377about 3 years ago
Dynamic Meta-Embeddings for Improved Sentence Representations
Pytorch332almost 6 years ago
DME/CDME332almost 6 years ago
Analogical Reasoning on Chinese Morphological and Semantic Relations
ChineseWordVectors11,874almost 3 years ago
Probabilistic FastText for Multi-Sense Word Embeddings
C++148over 8 years ago
Probabilistic FastText148over 8 years ago
Incorporating Syntactic and Semantic Information in Word Embeddings using Graph Convolutional Networks
TF291over 3 years ago
SynGCN
FRAGE: Frequency-Agnostic Word Representation
Pytorch118over 7 years ago
Wikipedia2Vec: An Optimized Tool for LearningEmbeddings of Words and Entities from Wikipedia
Cython946over 2 years ago
Wikipedia2Vec
Directional Skip-Gram: Explicitly Distinguishing Left and Right Context for Word Embeddings
ChineseEmbedding
cw2vec: Learning Chinese Word Embeddings with Stroke n-gram Information
C++274over 3 years ago
VCWE: Visual Character-Enhanced Word Embeddings
Pytorch15about 7 years ago
VCWE15about 7 years ago
Learning Cross-lingual Embeddings from Twitter via Distant Supervision
Text14over 6 years ago
An Unsupervised Character-Aware Neural Approach to Word and Context Representation Learning
TF0almost 8 years ago
ViCo: Word Embeddings from Visual Co-occurrences
Pytorch25about 7 years ago
ViCo25about 7 years ago
Spherical Text Embedding
C175almost 3 years ago
Unsupervised word embeddings capture latent knowledge from materials science literature
Gensim624over 3 years ago

awesome-sentence-embedding / OOV Handling

ALaCarte104almost 8 years ago:
Mimick153almost 7 years ago:
CompactReconstruction9over 3 years ago:

awesome-sentence-embedding / Contextualized Word Embeddings

Language Models are Unsupervised Multitask Learners
TF22,644about 2 years ago
117M22,644about 2 years agoGPT-2( , , , , , )
Learned in Translation: Contextualized Word Vectors
Pytorch473over 4 years ago
CoVe473over 4 years ago
Universal Language Model Fine-tuning for Text Classification
Pytorch26,390almost 2 years ago
EnglishULMFit( , )
Deep contextualized word representations
Pytorch11,774almost 4 years ago
AllenNLPELMO( , )
Efficient Contextualized Representation:Language Model Pruning for Sequence Labeling
Pytorch147over 6 years ago
LD-Net147over 6 years ago
Towards Better UD Parsing: Deep Contextualized Word Embeddings, Ensemble, and Treebank Concatenation
Pytorch1,462over 5 years ago
ELMo1,462over 5 years ago
Direct Output Connection for a High-Rank Language Model
Pytorch12over 7 years ago
DOC
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
TF38,374about 2 years ago
BERT38,374about 2 years agoBERT( , , )
Contextual String Embeddings for Sequence Labeling
Pytorch13,990almost 2 years ago
Flair13,990almost 2 years ago
Improving Language Understanding by Generative Pre-Training
TF2,167over 7 years ago
GPT2,167over 7 years ago
Multi-Task Deep Neural Networks for Natural Language Understanding
Pytorch2,238over 2 years ago
MT-DNN2,238over 2 years ago
BioBERT: pre-trained biomedical language representation model for biomedical text mining
TF1,970about 3 years ago
BioBERT672over 6 years ago
Cross-lingual Language Model Pretraining
Pytorch2,893over 3 years ago
XLM2,893over 3 years ago
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
TF3,619almost 4 years ago
Transformer-XL3,619almost 4 years ago
Efficient Contextual Representation Learning Without Softmax Layer
Pytorch4about 6 years ago
SciBERT: Pretrained Contextualized Embeddings for Scientific Text
Pytorch, TF1,532over 4 years ago
SciBERT1,532over 4 years ago
Publicly Available Clinical BERT Embeddings
Text680about 6 years ago
clinicalBERT
ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission
Pytorch386almost 4 years ago
ClinicalBERT
ERNIE: Enhanced Language Representation with Informative Entities
Pytorch1,413over 2 years ago
ERNIE
Unified Language Model Pre-training for Natural Language Understanding and Generation
Pytorch20,400almost 2 years ago
unilm1-large-casedUniLMv1( , )
HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document Summarization
Pre-Training with Whole Word Masking for Chinese BERT
Pytorch, TF9,746about 3 years ago
BERT-wwm9,746about 3 years ago
XLNet: Generalized Autoregressive Pretraining for Language Understanding
TF6,183over 3 years ago
XLNet6,183over 3 years ago
ERNIE 2.0: A Continual Pre-training Framework for Language Understanding
PaddlePaddle6,331about 2 years ago
ERNIE 2.06,331about 2 years ago
SpanBERT: Improving Pre-training by Representing and Predicting Spans
Pytorch893about 3 years ago
SpanBERT893about 3 years ago
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Pytorch30,675almost 2 years ago
RoBERTa30,675almost 2 years ago
Subword ELMo
Pytorch12over 6 years ago
Knowledge Enhanced Contextual Word Representations
TinyBERT: Distilling BERT for Natural Language Understanding
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Pytorch10,804almost 2 years ago
BERT-345MMegatron-LM( , )
MultiFiT: Efficient Multi-lingual Language Model Fine-tuning
Pytorch284over 6 years ago
Extreme Language Model Compression with Optimal Subwords and Shared Projections
MULE: Multimodal Universal Language Embedding
Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks
K-BERT: Enabling Language Representation with Knowledge Graph
UNITER: Learning UNiversal Image-TExt Representations
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
TF3,942almost 4 years ago
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Pytorch30,675almost 2 years ago
bart.baseBART( , , , , )
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Pytorch, TF2.0136,357almost 2 years ago
DistilBERT136,357almost 2 years ago
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
TF6,215almost 2 years ago
T56,215almost 2 years ago
CamemBERT: a Tasty French Language Model
CamemBERT
ZEN: Pre-training Chinese Text Encoder Enhanced by N-gram Representations
Pytorch645about 4 years ago
Unsupervised Cross-lingual Representation Learning at Scale
Pytorch2,893over 3 years ago
xlmr.largeXLM-R (XLM-RoBERTa)( , )
ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
Pytorch694about 2 years ago
ProphetNet-large-16GBProphetNet( , )
CodeBERT: A Pre-Trained Model for Programming and Natural Languages
Pytorch2,281about 3 years ago
CodeBERT
UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
Pytorch20,400almost 2 years ago
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
TF2,342over 2 years ago
ELECTRA-SmallELECTRA( , , )
MPNet: Masked and Permuted Pre-training for Language Understanding
Pytorch288about 5 years ago
MPNet
ParsBERT: Transformer-based Model for Persian Language Understanding
Pytorch341over 3 years ago
ParsBERT
Language Models are Few-Shot Learners
InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training
Pytorch20,400almost 2 years ago

awesome-sentence-embedding / Pooling Methods

SIF1,084about 7 years ago:
TF-IDF9over 3 years ago:
P-norm186over 5 years ago:
DisC54over 6 years ago:
GEM19over 7 years ago:
SWEM284almost 4 years ago:
VLAWE10over 7 years ago:
Efficient Sentence Embedding using Discrete Cosine Transform
fse: Gensim add-on for fast sentence embeddings. Supports Mean, Max, SIF, uSIF618over 3 years ago
Efficient Sentence Embedding via Semantic Subspace Analysis

awesome-sentence-embedding / Encoders

Incremental Domain Adaptation for Neural Machine Translation in Low-Resource Settings
Python5about 7 years ago
Distributed Representations of Sentences and Documents
Pytorch413almost 4 years ago
Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
Theano427over 9 years ago
Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
Theano2,050over 6 years ago
Order-Embeddings of Images and Language
Theano186almost 10 years ago
Towards Universal Paraphrastic Sentence Embeddings
Theano193over 10 years ago
From Word Embeddings to Document Distances
C, Python538over 2 years ago
Learning Distributed Representations of Sentences from Unlabelled Data
Python124over 9 years ago
Charagram: Embedding Words and Sentences via Character n-grams
Theano125about 10 years ago
Learning Generic Sentence Representations Using Convolutional Neural Networks
Theano34about 9 years ago
Unsupervised Learning of Sentence Embeddings using Compositional n-Gram Features
C++1,194about 4 years ago
Learning to Generate Reviews and Discovering Sentiment
TF1,512about 3 years ago
Revisiting Recurrent Networks for Paraphrastic Sentence Embeddings
Theano33over 9 years ago
Supervised Learning of Universal Sentence Representations from Natural Language Inference Data
Pytorch2,282about 5 years ago
VSE++: Improving Visual-Semantic Embeddings with Hard Negatives
Pytorch492almost 5 years ago
Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm
Keras1,525about 2 years ago
StarSpace: Embed All The Things!
C++3,948almost 4 years ago
DisSent: Learning Sentence Representations from Explicit Discourse Relations
Pytorch33over 6 years ago
Pushing the Limits of Paraphrastic Sentence Embeddings with Millions of Machine Translations
Theano102almost 3 years ago
Dual-Path Convolutional Image-Text Embedding with Instance Loss
Matlab287over 3 years ago
An efficient framework for learning sentence representations
TF205about 7 years ago
Universal Sentence Encoder
TF-Hub
End-Task Oriented Textual Entailment via Deep Explorations of Inter-Sentence Interactions
Theano16over 8 years ago
Learning general purpose distributed sentence representations via large scale multi-task learning
Pytorch311about 6 years ago
Embedding Text in Hyperbolic Spaces
TF8about 9 years ago
Representation Learning with Contrastive Predictive Coding
Keras527about 7 years ago
Context Mover’s Distance & Barycenters: Optimal transport of contexts for building representations
Python21over 5 years ago
Learning Universal Sentence Representations with Mean-Max Attention Autoencoder
TF16almost 8 years ago
Learning Cross-Lingual Sentence Representations via a Multi-task Dual-Encoder Model
TF-Hub
Improving Sentence Representations with Consensus Maximisation
BioSentVec: creating sentence embeddings for biomedical texts
Python578about 3 years ago
Word Mover's Embedding: From Word2Vec to Document Embedding
C, Python81almost 8 years ago
A Hierarchical Multi-task Approach for Learning Embeddings from Semantic Tasks
Pytorch1,191about 3 years ago
Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
Pytorch3,604over 2 years ago
Convolutional Neural Network for Universal Sentence Embeddings
Theano2over 8 years ago
No Training Required: Exploring Random Encoders for Sentence Classification
Pytorch184over 6 years ago
CBOW Is Not All You Need: Combining CBOW with the Compositional Matrix Space Model
Pytorch21over 7 years ago
GLOSS: Generative Latent Optimization of Sentence Representations
Multilingual Universal Sentence Encoder
TF-Hub
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Pytorch15,556almost 2 years ago
SBERT-WK: A Sentence Embedding Method By Dissecting BERT-based Word Models
Pytorch178over 5 years ago
DeCLUTR: Deep Contrastive Learning for Unsupervised Textual Representations
Pytorch380over 3 years ago
Language-agnostic BERT Sentence Embedding
TF-Hub
On the Sentence Embeddings from Pre-trained Language Models
TF530over 5 years ago

awesome-sentence-embedding / Evaluation

decaNLP2,345over 2 years ago:
SentEval2,086over 2 years ago:
GLUE779about 5 years ago:
Exploring Semantic Properties of Sentence Embeddings
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Word Embeddings Benchmarks437over 5 years ago:
MLDoc152over 4 years ago:
LexNET77,258almost 2 years ago:
wordvectors.net120over 5 years ago:
jiant1,650about 3 years ago:
jiant1,650about 3 years ago:
Evaluation of sentence embeddings in downstream and linguistic probing tasks
QVEC75over 8 years ago:
Grammatical Analysis of Pretrained Sentence Encoders with Acceptability Judgments
EQUATE : A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language Inference
Evaluating Word Embedding Models: Methods andExperimental Results
How to (Properly) Evaluate Cross-Lingual Word Embeddings: On Strong Baselines, Comparative Analyses, and Some Misconceptions
Linguistic Knowledge and Transferability of Contextual Representations:
LINSPECTOR24over 6 years ago:
Pitfalls in the Evaluation of Sentence Embeddings
Probing Multilingual Sentence Representations With X-Probe:

awesome-sentence-embedding / Misc

Word Embedding Dimensionality Selection329over 6 years ago:
Half-Size129over 5 years ago:
magnitude1,635about 3 years ago:
To Tune or Not to Tune? Adapting Pretrained Representations to Diverse Tasks
Don't Settle for Average, Go for the Max: Fuzzy Sets and Max-Pooled Word Vectors:
The Pupil Has Become the Master: Teacher-Student Model-BasedWord Embedding Distillation with Ensemble Learning:
Improving Distributional Similarity with Lessons Learned from Word Embeddings:
Misspelling Oblivious Word Embeddings:
Single Training Dimension Selection for Word Embedding with PCA
Compressing Word Embeddings via Deep Compositional Code Learning:
UER: An Open-Source Toolkit for Pre-training Models:
Situating Sentence Embedders with Nearest Neighbor Overlap
German BERT

awesome-sentence-embedding / Vector Mapping

Cross-lingual Word Vectors Projection Using CCA56about 8 years ago:
vecmap648over 3 years ago:
MUSE3,193about 4 years ago:
CrossLingualELMo99over 6 years ago:

awesome-sentence-embedding / Articles

Comparing Sentence Similarity Methods
The Current Best of Universal Word Embeddings and Sentence Embeddings
On sentence representations, pt. 1: what can you fit into a single #$!%@*&% blog post?
Deep-learning-free Text and Sentence Embedding, Part 1
Deep-learning-free Text and Sentence Embedding, Part 2
An Overview of Sentence Embedding Methods
Word embeddings in 2017: Trends and future directions
A Walkthrough of InferSent – Supervised Learning of Sentence Embeddings
A survey of cross-lingual word embedding models
Introducing state of the art text classification with universal language models
Document Embedding Techniques

Backlinks from these awesome lists: