ZEN
by sinovation
A BERT-based Chinese Text Encoder Enhanced by N-gram Representations
AI summary
Chinese text encoder
A pre-trained BERT-based Chinese text encoder with enhanced N-gram representations
- stars
- 645
- forks
- 104
- watching
- 22
Similar projects
Found by comparing what the projects do, not just their names.
Sentence encoder
A software package providing tools to encode and process text data using a specific neural network architecture.
Chinese character recognition network
This project demonstrates how to build and train a convolutional neural network (CNN) to recognize Chinese characters.
Chinese language models
Provides pre-trained models for Chinese language tasks with improved performance and smaller model sizes compared to existing models.
Token representation refinement
Improves pre-trained language models by encouraging an isotropic and discriminative distribution of token representations.
Handwritten Chinese recognizer
A Python-based web application that recognizes handwritten Chinese characters using a Convolutional Neural Network (CNN), allowing users to input text via an online writing board and receive recognition results.
Word embedding model
A Python implementation of a neural network model for learning word embeddings from text data
Chinese Tokenizer Library
A tokenizer based on dictionary and Bigram language models for text segmentation in Chinese
Chinese model
An implementation of a Chinese language model using PyTorch and transformer architecture.
Text transcription tool
Converts text to phonetic transcriptions in multiple languages using various backends and algorithms
Pinyin translator
A Python library for translating Chinese characters to pinyin
Visual Text Understanding
This project enables multi-modal language models to understand and generate text about visual content using referential comprehension.
Sentence encoder
Trains an autoencoder to learn generic sentence representations using convolutional neural networks
Chinese language model
Trains a large Chinese language model on massive data and provides a pre-trained model for downstream tasks
Word-based Chinese Model
A Word-based Chinese BERT model trained on large-scale text data using pre-trained models as a foundation
Chinese NLP Models
Provides pre-trained models for Chinese natural language processing tasks using the XLNet architecture