ZEN

Chinese text encoder

A pre-trained BERT-based Chinese text encoder with enhanced N-gram representations

A BERT-based Chinese Text Encoder Enhanced by N-gram Representations

GitHub

645 stars
22 watching
104 forks
Language: Python
last commit: about 4 years ago

Related projects:

RepositoryDescriptionStars
zminghua/sentencodingA software package providing tools to encode and process text data using a specific neural network architecture.16
soloice/chinese-character-recognitionThis project demonstrates how to build and train a convolutional neural network (CNN) to recognize Chinese characters.200
cluebenchmark/cluepretrainedmodelsProvides pre-trained models for Chinese language tasks with improved performance and smaller model sizes compared to existing models.806
yxuansu/taclImproves pre-trained language models by encouraging an isotropic and discriminative distribution of token representations.92
taosir/cnn_handwritten_chinese_recognitionA Python-based web application that recognizes handwritten Chinese characters using a Convolutional Neural Network (CNN), allowing users to input text via an online writing board and receive recognition results.511
zhangxiann/skip-gramA Python implementation of a neural network model for learning word embeddings from text data6
xujiajun/gotokenizerA tokenizer based on dictionary and Bigram language models for text segmentation in Chinese21
lonepatient/nezha_chinese_pytorchAn implementation of a Chinese language model using PyTorch and transformer architecture.262
bootphon/phonemizerConverts text to phonetic transcriptions in multiple languages using various backends and algorithms1,249
lxneng/xpinyinA Python library for translating Chinese characters to pinyin826
sy-xuan/pinkThis project enables multi-modal language models to understand and generate text about visual content using referential comprehension.79
zhegan27/convsentTrains an autoencoder to learn generic sentence representations using convolutional neural networks34
brightmart/xlnet_zhTrains a large Chinese language model on massive data and provides a pre-trained model for downstream tasks230
zhuiyitechnology/wobertA Word-based Chinese BERT model trained on large-scale text data using pre-trained models as a foundation460
ymcui/chinese-xlnetProvides pre-trained models for Chinese natural language processing tasks using the XLNet architecture1,652