tiny_segmenter
by 6
Ruby port of TinySegmenter.js for tokenizing Japanese text
AI summary
Text tokenizer
A Ruby port of a Japanese text tokenization algorithm
- stars
- 21
- forks
- 1
- watching
- 3
Similar projects
Found by comparing what the projects do, not just their names.
Sentence tokenizer
A Ruby port of the NLTK algorithm to detect sentence boundaries in unstructured text
Text tokenizer library
A Ruby-based library for splitting written text into tokens for natural language processing tasks.
Tokenizer
A gem for extracting words from text with customizable tokenization rules
Tokenizer
A Ruby library that tokenizes input and provides various statistical measures about the tokens
Chinese tokenizer library
Provides a Ruby port of the popular Chinese language processing library Jieba
Sentence tokenizer
A Ruby library that tokenizes text into sentences using a Bayesian statistical model
Text segmenter
Breaks text into contiguous sequences of words or phrases
Text parser
A simple tokenizer library for parsing and analyzing text input in various formats.
Tokenizer
A multilingual tokenizer to split strings into tokens, handling various language and formatting nuances.
Word tokenizer
A Python wrapper around the Thai word segmentator LexTo, allowing developers to easily integrate it into their applications.
Tokenizer
A fast and simple tokenizer for multiple languages
Language tagger
A Ruby wrapper for a statistical language modeling tool for part-of-speech tagging and chunking
Chinese Tokenizer Library
A tokenizer based on dictionary and Bigram language models for text segmentation in Chinese
Text tokenizer
A tool for tokenizing raw text into words and sentences in multiple languages, including Hungarian.
ismasan/oat278
Serializer
Provides a standardized way to serialize API data in Ruby apps.