Awesome Lists
C++pushed almost 2 years ago

Unicode tokeniser. Ucto tokenizes text files: it separates words from punctuation, and splits sentences. It offers several other basic preprocessing steps such as changing case that you can all use to make your text suited for further processing such as indexing, part-of-speech tagging, or machine translation. Ucto comes with tokenisation rules for several languages and can be easily extended to suit other languages. It has been incorporated for tokenizing Dutch text in Frog, our Dutch morpho-syntactic processor. http://ilk.uvt.nl/ucto --

AI summary

Text tokenizer

A tokeniser for natural language text that separates words from punctuation and supports basic preprocessing steps such as case changing

stars
66
forks
13
watching
13
awesome lists
2
View on GitHublanguagemachines.github.io/ucto

Embed the badge

Show how many awesome lists link to your project. The count updates automatically.

Awesome Lists badge
Markdown
[![Awesome Lists Badge](https://awesome.facts.dev/shield/LanguageMachines/ucto/links.svg)](https://awesome.facts.dev/awesome/LanguageMachines/ucto)
HTML
<a href="https://awesome.facts.dev/awesome/LanguageMachines/ucto"><img src="https://awesome.facts.dev/shield/LanguageMachines/ucto/links.svg" alt="Awesome Lists Badge" /></a>
Image URL
https://awesome.facts.dev/shield/LanguageMachines/ucto/links.svg

Add a GitHub project

Missing a project or an awesome list? Paste its GitHub URL and we fetch it right away.