jieba

Chinese tokenizer

A comprehensive Python library for Chinese text segmentation and word extraction.

结巴中文分词

GitHub

33k stars
1k watching
7k forks
Language: Python
last commit: about 2 years ago
Linked from 5 awesome lists


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
452896915/jieba-androidAn Android implementation of the Chinese word segmentation algorithm jieba, optimized for fast initialization and tokenization153
fukuball/jieba-phpA PHP module for Chinese text segmentation and word breaking1,331
mimosa/jieba-jrubyProvides a Ruby port of the popular Chinese language processing library Jieba8
mxrch/penglabA Google Colab setup for cracking hashes using multiple tools929
mmmaaaggg/ibats_huobifeeder_oldAutomates real-time market data retrieval and storage from Huobi exchange, publishing updates to Redis for use in backtesting and analysis.39
ioseb/geokbdA JavaScript library designed to simplify Georgian keyboard layout support57
toshi0383/ipanemaAnalyzes and prints useful information from IPA files used in iOS app development.10
edolphin-ydf/goimpl.nvimGenerates stubs for interfaces in code completion tools60
jalkoby/squasherA tool to compress and remove unnecessary migration history from database schema1,499
zerbea/hcxtoolsConverts packet capture files to usable hashes for Hashcat or John the Ripper analysis.2,039
hustcc/babel-plugin-optimize-i18nOptimizes internationalization text files by reducing bundle size through code substitution14
ma-ha/kicad-laser-stencil-pluginGenerates G-Code files for laser cutting solder paste stencils in KiCAD PCBs.16
rek7/mxtractAnalyzes and dumps memory to extract sensitive information from running processes582
reb311ion/replicaAn enhancement tool for Ghidra's binary analysis capabilities289
hobbyquaker/hm-discoverA tool to scan and discover Homematic devices on a network.6