low-resource-languages

Language preservation toolkit

A repository of tools and resources to support the documentation, conservation, and development of endangered languages.

Resources for conservation, development, and documentation of low resource (human) languages.

GitHub

393 stars
35 watching
56 forks
Language: TeX
last commit: over 2 years ago
Linked from 3 awesome lists

awesomeawesome-listendangered-languageshuman-languagelanguage-documentationlanguage-learninglanguage-resourceslistlow-resource-languageslrlsminority-languagenatural-languagenatural-language-processingnlpresourced-languages

Generic Repositories / Single language lexicography projects and utilities / Utilities

Project for Free Electronic DictionariesIs a project for a java MIDlet for mobile phones - for indigenous language dictionaries
WebonarySite which hosts digital dictionaries for single languages
WeSay18almost 2 years agoAllows language communities to build their own dictionaries. (by the SIL International)

Generic Repositories / Software

4lang37over 2 years agoConcept dictionary using Eilenberg machines
accentuate.usa.k.a. "charlifter". Statistical Unicodification of plain text for many languages
alignment-with-openfst21almost 10 years agoThis is an implementation of the CRF autoencoder framework for four tasks: bitext word alignment, part-of-speech tagging, code switching, dependency parsing
ApertiumApertium is a toolbox to build open-source shallow-transfer machine translation systems, especially suitable for related language pairs: it includes the engine, maintenance tools, and open linguistic data for several language pairs
ark-tweet-nlp0about 14 years agoCMU ARK Twitter Part-of-Speech Tagger ( )
ArtOfReading1over 6 years agoIndex and processing scripts related to the Art Of Reading illustration collection
bayesline0over 9 years agoA Multinomial Bayesian Classification for Language Identification
bible-corpus-tools15almost 4 years agoA collection of tools for reading/processing the multilingual Bible corpus
BloomDesktop39almost 2 years agoBloom Desktop is a hybrid c#/javascript/html/css Windows application that dramatically "lowers the bar" for language communities who want books in their own languages. Bloom delivers a low-training, high-output system where mother tongue speakers and their advocates work together to foster both community authorship and access to external materia…
BloomLibrary4over 5 years agoBloom Library Single Page App, using AngularJS & Bootstrap, Parse.com backend.
brain1over 12 years agoNeural networks in JavaScript
Bristol Uni MT Morphology tools2almost 11 years agoThis repo is a mirror of scripts previously available on . Included: Ukwabelana - An open-source morphological Zulu corpus and EMMA: A Novel Evaluation Metric for Morphological Analysis
brown-cluster425about 3 years agoC++ implementation of the Brown word clustering algorithm
CasualConCasualConc is a concordance program that runs natively on Mac OS X 10.5 Leopard or later. It was originally designed for casual use (preliminary analysis or non-research purposes), though [the maintainer] has been using it for his own research (and may others have). It can generate kwic concordance lines, word clusters, collocation analysis, and word count
cdec183over 6 years agoDecoder, aligner, and model optimizer for statistical machine translation and other structured prediction models based on (mostly) context-free formalisms
charlintCharlint is a character normalization/checking tool written in Perl. Among else, it implements Normalization Form C of Unicode TR 15, as a test platform for Early Uniform Normalization in the W3C Character Model
chorus7almost 2 years agoA version control system designed to enable workflows appropriate for typical language development teams who are geographically distributed
clam130over 2 years agoComputational Linguistics Application Mediator -- Quickly turn NLP applications into RESTful webservices with a web-application front-end. You provide a specification of your command line application, its input, output and parameters, and CLAM wraps around your application to form a fully fledged RESTful webservice
CMU SphinxCMUSphinx is a speaker-independent large vocabulary continuous speech recognizer released under BSD style license. It is also a collection of open source tools and resources that allows researchers and developers to build speech recognition systems
cnminlangwebcollect1almost 6 years agoChinese minorities website languages detection and websites collection
Cog23almost 3 years agoCog is a tool for comparing languages using lexicostatistics and comparative linguistics techniques. It can be used to automate much of the process of comparing word lists from different language varieties.
convertextract11about 3 years agoConvert Excel, Word and PowerPoint files with non-Unicode text (like text requiring SIL fonts) into Unicode, while preserving original file's formatting
CorpusTools115almost 2 years agoPhonological CorpusTools
CTK18over 10 years agoBuilt around LDC's champollion sentence aligner kernel, Champollion Tool Kit (CTK) aims to providing ready-to-use parallel text sentence alignment tools for as many language pairs as possible. (Original project is on SourceForge: )
DataTags0almost 12 years agoA system to assess the sensitivity and privacy risk of a dataset, and assign a tag to describe how the dataset must be transfered, stored and accessed. ( )
dataverse894almost 2 years agoA data repository framework to share and publish research data
Dative14over 3 years agoDative: software for linguistic fieldwork
dative14over 3 years agoA single-page application that interacts with multiple linguistic fieldwork web service databases.
DeepLearnToolbox0over 12 years agoMatlab/Octave toolbox for deep learning. Includes Deep Belief Nets, Stacked Autoencoders, Convolutional Neural Nets, Convolutional Autoencoders and vanilla Neural Nets. Each method has examples to get you started
Desmeme4over 2 years agoDatabase and tools for exploring linguistic templates
dictdbdictionary database for language translation
discoursegraphs50over 3 years agoPython-based tool to convert and merge multilayer annotated linguistic data
divvun-gramcheck9almost 2 years agoThis program does FST lookup on forms specified as Constraint Grammar format readings, and looks up error-tags in an XML file with human-readable messages. It is meant to be used as a late stage of a grammar checker pipeline
divvun-keyboard6almost 2 years agokeyboard apps for iOS and Android with keyboard layouts for indigenous and minority languages
divvunspell14almost 2 years ago(below) rewritten in Rust, for robust concurrency and memory management. Is in practical use about 10x faster than . It uses the same zhfst files as , which are available for all languages in the GitHub org (see below)
DLTK12about 11 years agoDeutsch Language Tool Kit.
epitran668about 2 years agoGrapheme to Phoneme conversion (G2P) for many low-resource languages
ELDER: Endangered Language Data Electronic Repository4almost 15 years agoEndangered Language Data Electronic Repository: A web-based ontologically-compliant collaborative linguistic data cataloguing tool
enchant1almost 2 years agoenchant spellchecking library
exsite97over 2 years agoExSite9 is a desktop application that was built to facilitate researchers easily and quickly tagging their data files with descriptive metadata and subsequently packaging their data files and associated metadata ready for submission to a repository. ExSite9 also allows for the structural organisation of said files within actually moving their physical location on your local file storage; allowing you to correctly organise your files and metadata ready for packaging
fast_align740about 4 years agoSimple, fast unsupervised word aligner
fastText25,979over 2 years agoLibrary for fast text representation and classification
FieldWorks86almost 2 years agoFieldWorks is a suite of software tools for language and cultural data, with support for complex scripts. FieldWorks Language Explorer (or FLEx, for short) is designed to help field linguists perform many common language documentation and analysis tasks. It can help you: elicit and record lexical information, create dictionaries, interlinearize texts, analyze discourse features, study morphology
Franc4,158over 2 years agoNatural language detection
FwDocumentation8almost 2 years agoDeveloper documentation for FieldWorks (software tools for language and cultural data, with support for complex scripts)
FwLocalizations0over 2 years agoLocalizations for FieldWorks
FwSupportTools2over 2 years agoAdditional tools for FieldWorks development
Gaia2,096about 5 years agoGaia is a HTML5-based Phone UI for the Boot 2 Gecko Project. NOTE: For details of what branches are used for what releases, see . If you're interested in setting up a keyboard in new language, see
giellakbd-android12about 2 years agoA fork of LatinIME (by Google for Android), targeting marginalised languages that also deserve first-class status on mobile operating systems. Used by (see elsewhere on this page)
giellakbd-ios30almost 2 years agoAn open source reimplementation of Apple's native iOS keyboard with a specific focus on support for localised keyboards. Used by (see elsewhere on this page)
giza-pp264over 3 years agoGIZA++ is a statistical machine translation toolkit that is used to train IBM Models 1-5 and an HMM word alignment model. This package also contains the source for the mkcls tool which generates the word classes necessary for training some of the alignment models
gv-crawl9almost 12 years agoGlobal Voices bitext crawler for creating parallel corpora
GlotLID106almost 2 years agoFasttext language identification with support for more than 2000 labels
Glottolog data12almost 9 years agoprovides comprehensive reference information for the world's languages
Gramadóir13about 3 years agoGrammar checking engine that is designed for the rapid development of grammar checkers for minority languages and other languages with limited computational resources
grind5about 6 years agoAn InDesign 5.5 plug-in designed allow graphite enabled smart fonts to be used in Adobe InDesign. This project integrates SIL's Graphite 2 smart font technology with our own implementation of a paragraph composer plugin
hermitcrab1over 4 years agoHermitCrab.NET is a flexible morphological/phonological parser that takes an item-and-process approach
hfst-ospell13over 2 years agoHFST spell checker library and command line tool
hfst-ospell-js0almost 10 years agoNode bindings for hfst-ospell
hfst-optimized-lookup12over 8 years agoHFST optimized-lookup standalone library and command line tool
hundict22about 12 years agobilingual dictionary extractor from parallel corpora
hunspell2,171almost 2 years agoSpell checker and morphological analyzer library and program designed for languages with rich morphology and complex word compounding or character encoding
huntag22over 10 years agoa sequential tagger for NLP using Maximum Entropy Learning and Hidden Markov Models
icu-dotnet62almost 2 years agoC# wrapper for ICU4C
icu4c6about 8 years agoMirror of svn project at . The FieldWorks branch has some FieldWorks specific enhancements
iLanguage21almost 9 years agoA semi-unsupervised language independent morphological analyzer useful for stemming unknown language text, or getting a rough estimate of possible parses for morphemes in a word. Input: a corpus. Uses compression, maximum entropy and fieldlinguistics
ipa-help0over 8 years agoIPA Helps
itweets-geodata0over 5 years agoGeodata from Indigenous Tweets
jQuery.ime175almost 2 years agojQuery based input methods library
kbdgen16over 2 years agoGenerate keyboards and keyboard layouts for various operating systems
koreksyon3about 11 years agoTools for developing and implementing spell-checking and grammar-checking capabilities in low-resource languages
l20n.js902over 7 years agoL20n reinvents software localization. Users should be able to benefit from the entire expressive power of natural languages. L20n keeps simple things simple, and at the same time makes complex things possible. This is the JavaScript implementation of L20n.
langid.py2,328over 6 years agoStand-alone language identification system
langtechA host of resources provided in SVN by the University of Tromsø. Details are and in English
LEGO Unified Concepticon0about 13 years agoMaterial relating to the LEGO Unified Concepticon
Lex4All21about 6 years agopronunciation LEXicons for Any Low-resource Language
lexdbLexDB is a lexical cognate tracking database. It stores the full provenance of all lexemes and cognate judgements, and allows export into a number of nexus dialects. The database is written in the flexible python/django web framework
LfMerge2almost 2 years agoSend/Receive for languageforge.org
liblevenshtein67almost 6 years agoA library for generating Finite State Transducers based on Levenshtein Automata
libpalaso44almost 2 years agoPalaso Library: A set of .Net libraries useful for developers of Language Software
LinGO Grammar MatrixThe LinGO Grammar Matrix is a framework for the development of broad-coverage, precision, implemented grammars for diverse languages
Lingpy126almost 3 years agoLingPy: Python library for quantitative tasks in historical linguistics
LinguisticaLinguistica is a program designed to explore the unsupervised learning of natural language, with primary focus on morphology (word-structure). It runs under Windows, Mac OS X and Linux, and is written in C++ within the Qt development framework. Its demands on memory depend on the size of the corpus analyzed
long-press305almost 7 years agojQuery plugin to ease the writing of accented or rare characters.
low-resource-pos-tagging-20149over 10 years agoLow-Resource POS-Tagging: 2014
lrl2over 13 years agoFor work concerning low resource languages
MacVoikko6over 11 years agoAn OS X spelling server based on Voikko
Machine28almost 2 years agoMachine is a natural language processing library for .NET that is focused on providing tools for processing resource-poor languages (used by FLEx)
Make-extensions6almost 9 years agoScripts for generating hunspell spellchecking extensions
mgiza161over 5 years agoA word alignment tool based on famous GIZA++, extended to support multi-threading, resume training and incremental training
Minority TranslateMinority Translate is a simple program for helping content generation on smaller sized Wikipedias (actually any sized) by giving pointers to existing articles in other language Wikipedias, so that the user can easily translate or adapt existing texts and thus increase the size and useability of their Wikipedia editions
morfessor186almost 6 years agoMorfessor is a tool for unsupervised and semi-supervised morphological segmentation
morpholm3about 13 years agoMorphology-aware language models
morph-test2over 5 years agoA python script to run tests for generation and analysis of a morphological transducer built using the Giella infrastructure. Works with Hfst, Xerox' fst tools, and with Foma
mosesdecoder1,585over 2 years agoMoses, the machine translation system
moz-l10n-tiers0almost 13 years agoCreates a pseudo-locale to evaluate string prioritization for l10n
mukurtucms84almost 2 years agoThe Mukurtu Content Management System (CMS) is an Internet- based platform designed to enable archiving of digital cultural resources
mythes40about 3 years agoMyThes is a simple thesaurus that uses a structured text data file and an index file with binary search to lookup words and phrases and return information on part of speech, meanings, and synonyms
myWorkSafe1about 8 years agoSmart & Simple Backup for Language Development Workers.
nabu19almost 2 years agonabu is a digital media item management system that provides a catalog of audio and video items, metadata for these items, and information about the workflow status of the items
Natural10,670about 2 years agogeneral natural language facilities for node
NIST 2008 Open Machine Translation Evalutation
NLTK13,694almost 2 years agoNatural Language Tool Kit. NLTK Source
node-panlex6over 7 years agonode.js client for PanLex
norma20over 5 years agoA tool for automatic spelling normalization
nplm14about 11 years agoFork of with some efficiency tweaks and adaptation for use in mosesdecoder
octothorpe0over 13 years agoCouchDB-powered wiki thing
OdtXslt2about 9 years agoPerform XSLT transform on contents of a package (such as ODT, Docx, etc.)
old-webapp4over 11 years agoOnline Linguistic Database --- software for creating web applications to collaboratively document languages.
old1about 6 years agoThe Online Linguistic Database (OLD): software for linguistic fieldwork.
old-pyramid8over 3 years agoOnline Linguistic Database migrated to the Pyramid framework
OmegaT-hfst-tokenizer2over 6 years agoOmegaT-hfst-tokenizer provides fst-based tokenisation in OmegaT
OpenDataKitOpen Data Kit (ODK) is an open-source suite of tools that helps organizations author, field, and manage mobile data collection solutions
OpenNLP1,449almost 2 years agoThe Apache OpenNLP library is a machine learning based toolkit for the processing of natural language text.
ops-devbox8over 3 years agoAnsible playbook for a (linux) developer machine
panlex-tools8almost 4 years agoThis package contains scripts to transform lexical resources into a format suitable for importing into PanLex. Documentation may be found at
pdsc-collection-viewer4almost 4 years agoParadisec Collection Browser
paradigm1almost 6 years agoPARADIGM is a .Net (C#) implementation of Joseph E. Grimes' 1983 work entitled "Affix Positions and Cooccurrences: The PARADIGM Program"
pathway7over 2 years agoPreparing language data for publication
pdfdroplet7almost 2 years agoLibrary and GUI for imposition of PDF pages (e.g. 2-up)
pepper23almost 2 years agoPepper is a pluggable, Java-based, open source converter framework for linguistic data
phonology-assistant10almost 4 years agoPhonology Assistant is a discovery tool. Provided with a corpus of phonetic data, it automatically charts the sounds and through its searching capabilities, helps a user discover and test the rules of sound in a language
pressagio19almost 7 years agoPressagio is a library that predicts text based on n-gram models. For example, you can send a string and the library will return the most likely word completions for the last token in the string
PrimerPro1almost 8 years agoThe purpose of PrimerPro is to assist the literacy worker in the development of primers for a given language
pyDelphin80about 2 years agoPython libraries for DELPH-IN (Friendly Fork)
RBGParser46over 10 years agoGraph-based Dependency Parser
Rosetta Pangloss0over 11 years agoThe Rosetta Project's Pangloss system
salm11over 8 years agoSALM: Suffix Array and its Applications in Empirical Language Processing by Joy
Salt15over 3 years agoA graph-based model to store and manipulate linguistic data
saymore6almost 2 years agoA tool for making common Language Documentation tasks such as keeping all the resulting files and meta data organized, converting files to archive formats, and transcription
Secwepemc-Facebook13over 11 years agoTranslate Facebook into unsupported languages
SegParser9almost 11 years agoRandomized Greedy algorithm for joint segmentation, POS tagging and dependency parsing
SeedLing11over 8 years agoBuilding and Using A Seed Corpus for the Human Language Project
Skype in your language3almost 11 years agoTranslate Skype into unsupported languages
solid1almost 2 years agoSolid is a software tool that can be used to check, clean up, and convert Standard Format (e.g. Toolbox) lexicon data
SPHERE Conversion ToolsMany LDC corpora contain speech files in NIST SPHERE format. The programs below convert SPHERE files to other formats
StandardFormatLib0over 11 years agoStandard Format Library
Stanford CoreNLP9,727almost 2 years agoStanford CoreNLP: A Java suite of core NLP tools.
Stanford CoreNLP Python612over 8 years agoPython wrapper for Stanford CoreNLP tools
stanza7,315almost 2 years agoStanford NLP group's shared Python tools
str2ipa10almost 11 years agoPronunciation dictionaries for languages with close-to-phonetic writing systems
sugali2about 4 years agoThis is a legacy repository of the language identification project for many (many) languages project for the software project course, NLP projects for low-resource languages
SuGarLike1about 12 years agoLanguage Identification for Low Resource Languages (by Susanne, Guy and Liling)
SyllabiPy44over 3 years agoPython interface for universal syllabification algorithms
tasty-imitation-keyboardA custom keyboard for iOS8+ that serves as a tasty imitation of the default Apple keyboard. Built using Swift and the latest Apple technologies!
TECkit18over 2 years agoA Text Encoding Conversion toolkit
teny3almost 14 years agoTools for low-resource machine translation
TeraDict6over 7 years agoTranslate English words into hundreds of languages!
Tesseract.js35,553almost 2 years agoPure Javascript OCR for 62 Languages 📖🎉🖥
TexNLP14over 14 years agoTexNLP: Texas Natural Language Processing tools
TiMBLTiMBL is an open source software package implementing several memory-based learning algorithms, among which IB1-IG, an implementation of k-nearest neighbor classification with feature weighting suitable for symbolic feature spaces, and IGTree, a decision-tree approximation of IB1-IG. All implemented algorithms have in common that they store some representation of the training set explicitly in memory. During testing, new cases are classified by extrapolation from the most similar stored cases
Toney5about 12 years agoTone Classification Software
Field Linguist's ToolboxToolbox is a data management and analysis tool for field linguists. It is especially useful for maintaining lexical data, and for parsing and interlinearizing text, but it can be used to manage virtually any kind of data
Toolbox Scripts for ELAN0over 11 years agoMirror of Alexander Koenig's Toolbox Scripts
ToolsForFieldLinguistics9over 7 years agoA collection of scripts and recipes for linguistics
transcriber2over 11 years agoAn HTML5 transcription tool for Aikuma
translitit-engine2over 8 years agoA transliteration engine written in JavaScript
Tsammalex data6about 8 years agois a multilingual lexical database on plants and animals
tweet2learn3over 7 years agoAn app to make it easier to use your native language on Twitter
twitter_langid15over 9 years agoA hierarchical character-word neural network for language identification
UniversalDependencies docs275almost 2 years agoUniversal Dependencies online documentation
UniversalDependencies tools207almost 2 years agoVarious utilities for processing the data
VocBenchVocBench is a web-based, multilingual, editing and workflow tool that manages thesauri, authority lists and glossaries using SKOS-XL
wavesurfer.js8,890almost 2 years agoNavigable waveform built on Web Audio and Canvas (Also has an ELAN plugin)
web-template3over 11 years agoThis is a web-based template that may be used to present language learning resources to aid language revitalization efforts. It includes a talking dictionary, and a phrasicon, containing sentences and phrases
webcorpus8over 11 years agoThis project is a collection of scripts and programs for creating a webcorpus from crawled data
wikt2dict53about 4 years agoWiktionary parser tool for many language editions
wikipron323almost 2 years ago-- retrives IPA pronunciations for Wiktionary entries
Word GeneratorWordGenerator generates hypothetical words from specifications of their syllable structure
WordBoundaryAn experiment in the detection and segmentation of word boundaries
wordbyword1about 12 years agoWordByWord is a free, open source, easy-to-use multimedia vocabulary trainer developed by Vera Ferreira, Peter Bouda, and Ricardo Filipe at CIDLeS with the support of the Foundation for Endangered Languages
WSI4URLang0almost 6 years agoWord Sense Induction (WSI) for Under-resourced Languages (URLang)
XDXF_Makedict228over 2 years agoXDXF dictionary format and "makedict" dictionary converting software (official repository)

Keyboard Layout Configuration Helpers

jQuery.IME175almost 2 years agojQuery Input Method Editor used on Wikipedia
kbdgen16over 2 years agoGenerate keyboards and keyboard layouts for Windows, macOS, X11, iOS, Android and Chrome, from a single, simple yaml file. Also registers languages unknown to Windows, so that after installation, there is a correct and robust association between the designated BCP 47 code (including full support for ISO 639-3) and installed language tools such as keyboards, spelling checkers and other tools
Keyboard1,780about 4 years agoVirtual Keyboard using jQuery ~
Keyboards153almost 2 years agoOpen Source Keyman keyboards
Keyman405almost 2 years agoKeyman cross platform input methods. Keyman makes it possible for you to type in over 1,000 languages on Windows, iPhone, iPad, Android tablets and phones, and even instantly in your web browser.
keyboardlayouteditor248over 4 years agoKeyboard Layout Editor
Keyboard layout editor1,332about 2 years agoKeyboard Layout Editor
lipika-ime117over 2 years agoInput Method Engine (IME) for Mac OS X with built-in support for all Indic Languages
XKeyboardConfigThe non-arch keyboard configuration database for X Window. The goal is to provide the consistent, well-structured, frequently released open source of X keyboard configuration data for X Window System implementations (free, open source and commercial). The project is targeted to XKB-based systems

Annotation

AGTK0over 10 years agoAGTK is a suite of software components for building tools for annotating linguistic signals, time-series data which documents any kind of linguistic behavior (e.g. audio, video). The internal data structures are based on annotation graphs. (Original project is on SourceForge: )
brendano8about 11 years agoGraph Fragment Language for Easy Syntactic Annotation
ELANELAN is a professional tool for the creation of complex annotations on video and audio resources
eopas9over 3 years agoETHNOER Online Presentation and Annotation System
FLAT - FoLia Linguistic Annotation Tool111about 2 years agoFLAT is a web-based linguistic annotation environment based around the FoLiA format ( ), a rich XML-based format for linguistic annotation. FLAT allows users to view annotated FoLiA documents and enrich these documents with new annotations, a wide variety of linguistic annotation types is supported through the FoLiA paradigm. It is a document-centric tool that fully preserves and visualises document structure
gfl_syntax8about 11 years agoGraph Fragment Language for Easy Syntactic Annotation
graf-python21about 12 years agoThe library graf-python is an open source Python implemenation to parse and write GrAF/XML files as described in ISO 24612. The parser of the library creates an annotation graph from the files. The user may then query the annotation graph via the API of graf-python
kwaras8almost 3 years agoTools for ELAN corpus management
LDC Word Aligner2over 8 years agoLDC Word Aligner is a software tool used for manual annotation of word alignment developed to support Arabic-English and Chinese-English word alignment tasks. It has a clean, easy-to-use interface. Since its development in 2009, LDC has used LDC Word Aligner to generate over 1,000,000 tokens of annotated word alignment data from a variety of genres including broadcast, newswire and web-based sources.
poio-analyzer13almost 13 years agoPoio is a collection of software tools for linguists working in language documentation, descriptive linguistics and/or language typology. It allows linguists to manage and analyze their data. The Poio Interlinear Editor allows to add morpho-syntactic annotations to transcriptions. It supports various file formats for input, but will only output standardized XML defined by the Corpus Encoding Standard and the Text Encoding Initiative. Several tools for analyzing linguistic data will be made available to further process annotated data. Poio tools are written in Python and are based on PyQt
poio-api18over 8 years agoPoio API is a free and open source Python library to access and search data from language documentation in your linguistic analysis workflow. It converts file formats like Elan’s EAF, Toolbox files, Typecraft XML and others into annotation graphs as defined in ISO 24612. Those graphs, for which we use an implementation called “Graph Annotation F…
pyannotation16about 14 years agoPyAnnotation is a Python Library to access and manipulate linguistically annotated corpus files
XTransTrans is a next generation multi-platform, multilingual, multi-channel transcription tool that supports manual transcription and annotation of audio recordings. The XTrans toolkit provides new and efficient solutions to common transcription challenges and addresses critical gaps in existing tools.Designed with input from experienced human transcribers working with real world data, XTrans provides a flexible and intuitive graphical user interface for a multitude of speech annotation tasks including (virtual) segmentation of audio into smaller units like turns and sentences; speaker identification; orthographic transcription in any language; and labeling of structural elements of the transcript like topics

Format Specifications

spec22over 3 years agoThe official specification for the DLx linguistic data format.
FoLiA61over 2 years agoFoLiA: Format for Linguistic Annotation - FoLiA is a rich XML-based annotation format for the representation of language resources (including corpora) with linguistic annotations. A wide variety of linguistic annotations are support, making FoLiA a useful format for NLP tasks and data interchange
xdxf_makedict228over 2 years agoXDXF dictionary format and "makedict" dictionary converting software (official repository)
Express-Lingua66over 12 years agoAn i18n middleware for the Express.js framework
Polyglot.jsGive your JavaScript the ability to speak many languages
TransifexSystem for providing a nice, userfriendly/project oriented approach to translating files. Great for non-technical users, free for open-source projects, decent for minority languages; , it can take a while to get a new language added to the Transifex system because the ticketing system Transifex uses results in them losing tickets sometimes. Provides translation memory, ability to appoint reviewers, etc. Transifex used to have an open source system that you could host on your own, but that seems to have disappeared

Audio automation

arctic-prompts1over 10 years agoGenerate prompts PDF for CMU ARCTIC dataset
AudioWebService4over 3 years agoa simple nodejs server which accepts upload of audio and runs it through praat
AuToBI58over 7 years agoAutomatic prosodic annotation tool written in Java
BashScriptsForPhonetics0almost 13 years ago( of a dormant project)
esv-text-audio-aligner93over 13 years agoESV Text/Audio Aligner to programmatically obtain the timings for each word in the corresponding audio
html5-audio-read-along192almost 9 years agoHTML5 Audio Read-Along
ipa-chart131over 5 years agoInternational Phonetic Alphabet (IPA) Unicode Chart and Character Picker
kaldi-svn-archive16about 11 years agoAn read-only archive of the original Kaldi SVN repository (mainly to keep sandboxes available)
lex4all1about 12 years agopronunciation LEXicons for Any Low-resource Language ( of a student project)
Montreal-Forced-Aligner1,364almost 2 years agoPython interface for forced text/speech alignment
node-pocketsphinx243over 7 years ago
opensauce5about 9 years agoGNU Octave-compatible version of VoiceSauce
pocketsphinx3,981almost 2 years agoPocketSphinx is a lightweight speech recognition engine, specifically tuned for handheld and mobile devices, though it works equally well on the desktop
pocketsphinx-ios-demo75about 8 years agoSimple demo for iOS
pocketsphinx-python338about 4 years agoPython module installed with setup.py
pocketsphinx-ruby13over 11 years agoRuby speech recognition with Pocketsphinx
pocketsphinx-wp-demo21over 10 years agoDemo to run pocketsphinx on WP8 platform
pocketsphinx.js1,493over 6 years agoSpeech recognition in JavaScript
praat-py0almost 14 years agoFrom my PhD days: Praat-Py is a custom build of Praat, the computer program used by linguists for doing phonetic analysis on sound files, to allow for scripts to be written in the Python programming language, rather than in Praat's built-in language. ( of a dormant project)
Praat-Scripts53almost 5 years agoMietta's Scripts
PraatTextGridJS12almost 5 years agoA small library which can parse TextGrid into json and json into TextGrid
PraatontheWeb39almost 5 years agoWeb implementation of Praat. Source code, running demo scripts on web, samples and documentation
prosodicParsing2over 14 years agodifferent kinds of HMMs to use for incorporating prosody into basic parsing
Prosodylab-Aligner333about 6 years agoPython interface for forced audio alignment using HTK and SoX
prosodylab.alignertools12over 11 years ago
Recordmp3js2about 11 years agoRecord MP3 files directly from the browser using JS and HTML
sphinx41,411almost 4 years agoPure Java speech recognition library
sphinxbase527over 4 years ago
sphinxtrain183almost 2 years ago
TLSphinx15over 7 years agoSwift wrapper around Pocketsphinx

Text-to-Speech (TTS)

espeakeSpeak is a compact open source software speech synthesizer for English and other languages, for Linux and Windows.
MARY TTS2,385almost 2 years agoMARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure java
OssianOssian is a collection of Python code for building text-to-speech (TTS) systems, with an emphasis on easing research into building TTS systems with minimal expert supervision

Automatic Speech Recognition (ASR)

Elpis152over 2 years agoElpis is software for creating speech recognition models and applying them to the transcription of audio. As of 2022, it gives access to Kaldi and Huggingface Transformers
kaldi14,362almost 2 years agoThis is now the official location of the Kaldi project
Persephone157over 3 years agoPersephone aims to make state-of-the-art phonemic transcription accessible to people involved in language documentation, who have a training corpus of about one to four hours of transcribed speech. As of 2022, Persephone is superseded by Elpis

Text automation

clld54almost 2 years agoCross Linguistic Linked Data python library
LaTeX2HTML561over 4 years agoLaTeX web components
MultilingualCorporaExtractor0over 13 years agoNode io Spider for extracting multilingual corpora ( of a student project)
SeedLing2about 12 years agoBuilding and Using A Seed Corpus for the Human Language Project ( of a student project)

Experimentation

experigen35about 6 years agoA framework for creating linguistic experiments
GamifyPsycholinguisticsExperiments0over 14 years agoA simple node server to gamify linguistics experiments, runs offline on a laptop for small scale experiements and online on a server for large scale experiments. Data is sent to a Google spreadsheet. ( of a dormant project)
OpenSesame242almost 2 years agoGraphical experiment builder for the social sciences
OPrime0almost 12 years agoOpen Source Experimentation Libraries - Online and Offline for Android and HTML5
psychopyMegProsody0over 13 years agoRuns MegProsody using PsychoPy
PsychScript4almost 12 years agoA HTML5/Javascript library for running behavioural experiments online

Flashcards

Anki19,289almost 2 years agoAnki is a program to make and share flaschard decks (including audio) for any language or writing system.
awesome-anki1,649almost 2 years agoA curated list of awesome Anki add-ons, decks and resources
VocabLift3about 12 years agoLanguage-learning tool that uses vocabulary from LIFT-format dictionaries produced by programs such as Fieldworks Language Explorer and WeSay

Natural language generation

OpenCCG206over 5 years agoOpenCCG library for parsing and realization with CCG. Includes mini-grammars for Inuit, Nezperce, Basque and others

Computing systems

Common Language Resources and Technology Infrastructure Norway / ClarinoOne of their projects (not clearly listed here) is about providing an online system for language analysis, so users can connect resources visually, dump in text, and get a result. Kind of like the Yahoo! Pipes but for language processing. Uses the cluster

Android Applications

Aikuma30over 10 years agoAndroid software for recording and translation
Android Speech Recognition Trainer3almost 8 years agoSpeech recognition training app for low resource languages which interfaces with FieldDB corpora
android-template0over 11 years agoThis is a template of an Android word-learning app that may be used a way to introduce a language. It includes a quiz. For the documentation, go to
AndroidFieldDB3over 7 years agoAn Android app which lets the user build a custom visual and auditory vocabulary, useful for guided anomia treatment and self designed language lessons by heritage speakers
AndroidFieldDBElicitationRecorder2almost 13 years agoA general purpose video recording tool
AndroidLanguageLessons2almost 8 years agoLets heritage speakers create self designed language lessons
AndroidProductionExperiment0almost 13 years agoAndroid App to run perception experiments
Bevara3almost 13 years agoAndroid Phone Application designed for Linguistic Fieldwork to help preserve, maintain, and save endangered languages
ojoVozA mobile app for sending georeferenced image and voice recordings from an Adroid phone to an email address. For more information, please go to
pocketsphinx-android235over 6 years agopocketsphinx build for Android
pocketsphinx-android-demo549almost 8 years ago

Chrome Extensions

babelfrog16over 7 years agoChrome extension to help learn languages as you browse
DictionaryChromeExtension6over 11 years agoDictionary for websites in low-resource languages. App and codebase which connects to a Wiktionary to provide definitions of any term on any website (current languages Cherokee 194,426 entries, Inuktitut 251 entries, Kartuli 7,363 entries, Plains Cree (incubation) 0 entries)

FieldDB

FieldDB79almost 4 years agoAn offline/online field database which adapts to its user's terminology and I-Language, has plugins for various data automation routines along the process of primary data collection to cleaning to publication and archival.

FieldDB / FieldDB Webservices/Components/Plugins

AndroidLanguageLearningClientForFieldDB-sikuli0almost 12 years agoSikuli tests for AndroidLanguageLearningClientForFieldDB
AuthenticationWebService0over 3 years agoA node.js web service which mananges users and corpora creation and authentication
bower-fielddb-angular0about 11 years agoA bower repository which hosts fielddb-angular components, bower install fielddb-angular --save
bower-fielddb0about 6 years agoA bower repository which hosts fielddb core components, bower install fielddb --save
fielddb-spreadsheet-sikuli1over 11 years agosikuli tests for the spreadsheet module
FieldDBActivityFeed0over 11 years agoA fielddb activity feed widget which can be embedded in other codebases, websites etc
FieldDBGlosser0over 9 years agoA semi-unsupervised language independent morphological analyzer useful for stemming unknown language text, or getting a rough estimate of possible parses for morphemes in a word. bower install fielddb-glosser --save
FieldDBLexicon0almost 9 years agoA lexicon browser/editor web widget for FieldDB databases
LanguageClassDashboard0about 12 years agoApp which provides a view of FieldDB corpora for language teachers
LexiconWebService0about 6 years agoA node.js ElasticSearch wrapper for indexing/training lexicons from corpora
LexiconWebServiceSample1over 14 years agoA node.js web server which implements the fieldlinguist's lexicon API for the FieldDB project

Academic Research Paper-Specific Repositories

Gargantua12almost 11 years agoFast Unsupervised Sentence Aligner described in "Improved unsupervised sentence alignment for symmetrical and asymmetrical parallel corpora", COLING 2010
ldc-kiy0about 13 years agoMaterials for: The experimental state of mind in elicitation: illustrations from tonal fieldwork. Dubmitted to Language Documentation & Conservation,
Learning to map into a Univerisal POS tagsetYuan Zhang, Roi Reichart, Regina Barzilay and Amir Globerson
low-resource-pos-tagging-20149over 10 years agoand Published in: Learning a Part-of-Speech Tagger from Two Hours of Annotation. . In Proceedings of NAACL 2013. And in: Real-World Semi-Supervised Learning of POS-Taggers for Low-Resource Languages. . In Proceedings of ACL 2013
orthotree10over 11 years agoLinguistic family tree based on orthographic distance
type-supervised-tagging-2012emnlp1over 10 years agoThis repository contains the code, scripts, and instructions needed to reproduce the results in the paper: Type-Supervised Hidden Markov Models for Part-of-Speech Tagging with Incomplete Tag Dictionaries. . In Proceedings of EMNLP 2012. This code is frozen as of the version used to obtain the results in the paper. It will not be maintained. To see the updated code, visit
visualizing-language1over 14 years agoFor visualizations of WALS and other typological databases
WALS-APiCS0over 11 years agoCode for working with WALS-APiCS (Atlas of Pidgin and Creole Language Structures) complexity metrics

Example Repositories

CorpusWebService0over 4 years agoüber-simple node.js-Proxy to enable CORS request for couchdb
CorporaForFieldLinguistics3about 9 years agoSmall corpora from diverse language typologies, useful for testing scripts
startR0almost 14 years ago
lucenerevolution-20130over 13 years agoDemo examples for linguistics in Lucene and Solr
berlin-buzzwords-20130over 13 years agoDemo examples for Lucene, Solr, ElasticSearch and OpenNLP from Berlin Buzzwords 2013 talk

Fonts

fontinline4about 8 years agoMake inline stroke paths from an outline font
Noto Fonts2,466over 3 years agoNoto is Google’s free font family that aims to support all the world’s scripts. Its design goal is to achieve visual harmonization across languages. Noto fonts are under Apache License 2.0
UnicodifyUnicodify is a suite of programs for converting text in a variety of 8-bit encodings to Unicode (using the UTF-16 encoding). Unicodify was particularly designed to handle HTML-based text using non-ISCII 8-bit fonts to render South Asian scripts. However, elements of the suite can map other types of non-ASCII 8-bit encodings, such as Latin-2, ISCII and PASCII

Corpora

bible-corpus177about 2 years agoA multilingual parallel corpus created from translations of the Bible
poio-corpus7almost 2 years agoThe Poio Corpus is a freely available collection of language resources for the lesser-used languages. The data is extracted from free sources like Wikipedia, dictionaries, documents, websites and others

Organizations / On GitHub

batumiSpeech recognition and natural language processing for low-resource languages
BloomBooks
unicode-cldrUnicode Common Locale Data Repository (CLDR) Project
cmusphinxMirror of the SourceForge repositories
dativebaseTools for working with OLD
divvunThe Divvun group at UiT develops proofing tools, keyboard apps and other language technology solutions for indigenous and minority languages, especially the Sámi languages.
FieldDB
GiellaLThome for keyboard layouts, lexicons and morphologies for indigenous and minority languages, especially for morphologically complex languages, using mainly rule-based techonlogies. The resources are used by Divvun (above) and Giellatekno (below) to build a number of tools for the language communities. Almost everything is open source
HFSTHelsinki Finite-State Technology.
hunspell
keymanapp
langtechLanguage Technology Group, University of Melbourne
lex4all
longnow
MontrealCorpusTools
moses-smtStatistical Machine Translation
mukurtucms
NLTKNatural Language Toolkit
PhonologicalCorpusTools)
Projet de recherche sur l'écritureCrowdsourcing or conducting large scale psycholinguistics experiments (or statistically significant field linguistics)
prosodylabProsodylab at McGill University, Canada
SIL International (Dev)Another SIL organization, with many repositories
SIL InternationalSIL (originally known as the Summer Institute of Linguistics, Inc.) is probably the leading organization which provides software and tools tailored for use by field linguists and lexicographers working on endangered languages. A little known fact is that much of it's code is open sourced on GitHub and SIL is happy to recieve open source contributions and collaborate on open source projects
SIL NRSISIL Non-Roman Script Initiative. The NRSI is a department of SIL International, whose task is to provide assistance, research and development for SIL International and its partners to support the use of non-Roman and complex scripts in language development
StanfordNLP
ucsd-field-labUniversity of California, San Diego
UniversalDependenciesUniversal Dependencies (UD) is a project that is developing cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on an evolution of (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008). The general philosophy is to provide a universal inventory of categories and guidelines to facilitate consistent annotation of similar constructions across languages, while allowing language-specific extensions when necessary
utcomplingThe University of Texas at Austin's Computational Linguistics Lab.

Organizations / Other OSS Organizations

GiellateknoGiellatekno combines cutting-edge linguistic and computational research into the analysis of Saami and other morphologically-rich languages, with the development of practical applications. We focus on deep linguistic modeling and on highly efficient and robust computational analysis with a wide empirical coverage. They use svn for their code: all of it can be found , sorted by language
LOWLANDSLOWLANDS – Parsing low-resource languages and domains
LTRC: Language Technologies Research Center IIIT HyderabadLTRC addresses the complex problem of understanding and processing natural languages in both speech and text mode. LTRC conducts research on both basic and applied aspects of language technology. It is the largest academic centre of speech and language technology in South Asia. LTRC carries out its work through four labs, which work in synergy with each other, as listed above
The Language ArchivePart of the MPI

Tutorials

How to Write a Spelling Correctorby

Language Specific Projects / Afrikaans

Afrikaanse rekenaarlinguïstiek (Afrikaans computational linguistics)— wordlists, corpora, morphological analyser, tagger, word decompounder. Available upon email

Language Specific Projects / Albanian

Apertium rules for AlbanianMachine Translation rules
out-of-copyright-albanian-authorsauthors scraped from the albanian language wikipedia who are out of copyright
Plis keyboardThe Plis keyboard is a keyboard or computer keyboard layout for the Albanian language
spell checkingHere you find a collection of Albanian words and information about them. Aspell, Ispell, and MySpell are included

Language Specific Projects / Alutiiq

wiinaq2over 3 years agoWord Wiinaq is a dictionary web application with automatically generated ending tables and souped-up search capabilities. It is written in Python using Django

Language Specific Projects / Amharic

HornMorpho5over 11 years agoMorphological analysis and generation of Amharic and Oromo verbs and nouns and Tigrinya verbs

Language Specific Projects / Basque

MatxinAn open-source transfer machine translation engine. Linguistic information for the translation from Spanish and Basque (es-eu) is included

Language Specific Projects / Bengali

Bangla-অঙ্কুর for MacThis project aims to develop a phonetic based Bangla typing system for Macintosh computer which can be developed into a transliteration technique in the future
Bengali Writer1over 10 years ago`Bengali Writer' is a set of utilities for computerized editing and typesetting in Bengali, a language of India and Bangladesh. It comprises a set of fonts for Bengali in several formats (METAFONT, BDF, PS), a text editor with spell-cheking, export, and more. (Original project is on SourceForge: )
EkusheyBangla Computing and Localization Project for the Bangla speaking people
Lekho0over 10 years agoA collection of tools and resources for using bangla on computers (Original project is on SourceForge: )

Language Specific Projects / Chichewa

Chichewa9over 5 years agoNLP resources for Chichewa

Language Specific Projects / Galician

an-metri-gal3almost 2 years agoAnálise métrico de texto en verso en lingua galega (Galician language) gl-ES
android_gl_dict2over 13 years agoAndroid Galician (gl_ES) Keyboard Dictionary
aspell-gl1about 14 years agoGalician dictionary for aspell
CitiusSentiment7over 10 years agoSentiment analysis (opinion mining) for Portuguese, English, Spanish, and Galician
CitiusTaggerA PoS-Tagger and Named Entity Classification tool for Portuguese, English, Galician, and Spanish
ConshugaGalician verb conjugator
corpora2over 10 years agoThis is a collection of corpus of Galician (or related to Galicia) words / Colección de corpus de palabras en galego (ou relacionadas con Galicia)
DepPattern10over 8 years agoDependency Syntactic Parsing for Portuguese, Spanish, English, and Galician, including MetaRomance parser
DOGA_scraper0about 12 years agoGalician Official journal scraper
elFinder-language1almost 10 years agoGalician - Gallego / language for elFinder
EuroWordNetLemon1about 11 years agoEuroWordNet lemon lexicons generated from the LMF versions of the Multilingual Central Repository (MCR) EuroWordNet lexicons. It includes lexicons for Spanish, Catalan, Basque & Galician
GalegoDroidGalician Translator for Android
galeXtra2over 10 years agoMultiword Extractor for Portuguese, English, Spanish, Galician, French
Galician-Dependency-Treebank1almost 10 years agoThis Galician Dependency Treebank has been developed by transliterating and adapting lexically the Portuguese part (Bosque 7.3 by the Floresta sintá(c)tica project) of the CONLL-X 2006
Galician-Fuzzy-Text-watch1over 10 years agoBased on Fuzzy Text International by Jesse Hallett, uses the galician language to display time
galician-locale-for-mac1over 10 years agoGalician locale for Mac OS X
gl-syllabler1over 10 years agoSplit galician language words into syllables
gl1about 2 years agoGalician OmegaT Localisation
hunspell-gl-ciencias0about 13 years agoProject oriented into developing a science and maths Galician language Hunspell dictionary
hunspell-gl1over 13 years agoGalician hunspell dictionaries
hyphen-gl1over 14 years agoGalician hyphenation rules
javagalician-java63over 14 years agoThe Java Galician Locale is an implementation of Java localization SPIs which will allow the Java VM to use the Galician Language (locales "gl" and "gl_ES"), one of the official languages of Spain, which is not included in Sun's JVM distribution
Linguakit65over 2 years agoMultilingual toolkit for NLP: dependency parser, PoS tagger, NERC, multiword extractor, sentiment analysis, etc
ParlamentoGalicia0over 13 years agoProject based on the information extracted from the transcriptions of the sessions held in the Galician Parlament
poss-gl1about 15 years agoGalician translation of Producing Open Source Software, by Karl Fogel
rima1over 10 years agoFind rhyming words in galician language
stopwords-gl1almost 10 years agoGalician stopwords collection
texlive-babel-galician1almost 2 years agoTeXLive babel-galician package
UD_Galician-CTG1almost 2 years agoThe Galician UD treebank is based on the automatic parsing of the Galician Technical Corpus created at the University of Vigo by the the TALG NLP research group
UD_Galician-TreeGal6almost 2 years agoThe Galician-TreeGal is a treebank for Galician developed at LyS Group (Universidade da Coruña)
UL_Galician-TreeGal0over 8 years agoCoNLL-UL Repository for UD_Galician-TreeGal

Language Specific Projects / Galician / Apertium

apertium-cat-glg1about 4 years agoApertium translation pair for Catalan and Galician
apertium-dict-en-gl1over 10 years agoEnglish-Galician language pair for Apertium
apertium-dict-es-gl1over 10 years agoSpanish-Galician language pair for Apertium
apertium-dict-pt-gl1about 13 years agoPortuguese-Galician language pair for Apertium
apertium-en-gl0over 4 years agoApertium translation pair for English and Galician
apertium-es-gl1about 5 years agoApertium translation pair for Spanish and Galician
apertium-glg0about 4 years agoApertium linguistic data for Galician
Apertium-pt-gl.pt-gl-LMF0about 12 years agoThis is the LMF version of the Apertium bilingual ditionary for Portugues and Galician languages
apertium-pt-gl0about 5 years agoApertium translation pair for Portuguese and Galician

Language Specific Projects / Georgian

awesome-georgia90about 3 years agoA curated list of awesome libraries and packages specific/related to Georgia (country)
Gadatsqvetilebebi1over 9 years agoგადაწყვეტილებები; Web spider and corpora importer for public legal decisions
GeoWordsDatabase70almost 9 years agoAround 310 000 unique Georgian words
Kartuli Speech Recognition4over 8 years agoანდროიდის ქართველი მომხმარებლებისთვის სიტყვის ამოცნობის სისტემის შექმნა. Codebase to turn any webpage from any alphabet into another alphabet, the default is to turn latin letters into Kartuli. "Do your friends keep commenting on Facebook with English keyboards (either because they forgot to switch, or because they didn't/can't install a Georgian keyboard)? Now you can read the web through კართული eyes."
KartuliChromeExtension1over 12 years agoChrome აპლიკაცია, რომელიც ყველა ინგლისურ ასო-ბგერას აჩვენებს ქართულ ასო-ბგერად
QartuliDaBunebismetkveleba1almost 13 years agoმათემატიკისა და ბუნებისმეტყველების ინტერაქტიული სახელმძღვანელო მე-2 - მე-3 კლასის მოსწავლეებისათვის
SakartvelosUzenaesiSasamartloSarke0over 12 years agoსაქართველოს უზენაესი სასამართლო სარკე
SamartlosSakonstitutsioSasamartdoSarke0over 9 years agoსამართლოს საკონსტიტუციო სასამართდო სარკე
translitit-latin-to-mkhedruli-georgian4over 9 years agoA Latin to ქართული (Mkhedruli Georgian) transliteration function written in JavaScript
translitit-mkhedruli-georgian-to-ipa0over 9 years agoA Latin to ქართული (Mkhedruli Georgian) transliteration function written in JavaScript
Declensions2almost 4 years agoMethods to generate declensions for Georgian language

Language Specific Projects / Georgian / Fonts

Stichoza/font-larisome39over 5 years agoIconic font for Georgian currency inspired by Font-Awesome (CSS)
Lotuashvili/BPGNateli0about 11 years agoBower package for BPG Nateli font (CSS)
thecotne/georgian-webfonts17about 9 years agoPackage for georgian fonts (CSS)

Language Specific Projects / Georgian / Internationalization and Localization (i18n/l10n)

Stichoza/money-num-to-string8over 2 years agoConvert a number/money to localized string (PHP, JavaScript)
natchkebiailia/NumberToWord3almost 9 years agoConvert numbers to localized strings (JavaScript)
d0ragon/number-to-words-ka3over 12 years agoConvert numbers to localized strings (PHP)
dimakura/ka0almost 13 years agoCommon functionality for georgian projects (Ruby)
dimakura/ka.js5about 12 years agoGeorgian language support for node and browser (JavaScript)
akalongman/kautilities4about 10 years agoConvert Georgian letters to Latin and vice-versa (PHP)
Landish/Laravel-Ka8over 5 years agoGeorgian Language Pack
Landish/RedactorJS-GERedactor WYSIWYG HTML Editor Georgian Language Pack (JavaScript)
wenzhixin/bootstrap-table11,746almost 2 years agoBootstrap table with extra features. l10n by and
moment/moment48,013about 2 years agoA lightweight date library (JavaScript)
ioseb/geokbd57almost 17 years agoGeorgian keyboard library (JavaScript)

Language Specific Projects / Guarani

ParaMorfo5over 11 years agomorphological analysis and generation of Spanish and Guarani verbs, nouns, and adjectives

Language Specific Projects / Hausa

Hausa6about 11 years agoRepository for Hausa NLP tools

Language Specific Projects / Hindi

hindi-morph0over 13 years agoAn open source morphological analyzer for Hindi

Language Specific Projects / Høgnorsk

hunspell-hn_NOA beginning to a spellchecking tool for Høgnorsk, a conservative variant of Norwegian Nynorsk, based on a set of corpuses

Language Specific Projects / Icelandic

IceNLP21over 2 years agoIceNLP is an open source Natural Language Processing (NLP) toolkit for analyzing and processing Icelandic text. The toolkit is implemented in Java

Language Specific Projects / Inuktitut

InuktitutAlignerData3over 14 years agoScripts for alignment of laboratory speech production data
InuktitutComputing10about 11 years agoInuktitut Morphological Analyser, transcoder, transliterator, corpus tools, and lexical lists for working with Inuktitut. Usable online at

Language Specific Projects / Irish

aimsigh1about 3 years agoSource for the now-defunct aimsigh.com Irish search engine
caighdean18about 2 years agoCode for standardizing Irish language text
fleiscin1almost 6 years agoIrish hyphenation patterns for TeX
GaelSpell17almost 2 years agoSources for an Irish language spell checker
tesseract-gle-uncial4over 11 years agoOCR for old Irish fonts

Language Specific Projects / Kinyarwanda

kin-morph-fst6about 13 years agoKinyarwanda morphological analyzer
TurboTagger & TurboParser for Kinyarwanda (download)TurboTagger & TurboParser for Kinyarwanda

Language Specific Projects / Kurdish

KurlexMorphological analyser and lexicon, written in the Alexina framework, licensed under the LGPL-LR
kurmanji-stemmer1about 11 years agoNLTK based kurmanji stemmer

Language Specific Projects / Lingala

Lingala NLPNLP tools and resources for Lingala

Language Specific Projects / Lushootseed

Lushootseed0over 10 years agoJoshua Crowgey's work on Lushootseed

Language Specific Projects / Malay

MorfoMalayu5over 11 years agomorphological analysis of Malay words

Language Specific Projects / Malagasy

Global Voices Malagasy ProjectThis page provides a link to a corpus of parallel news articles in Malagasy and English from the Global Voices project. This corpus was collected and aligned at the sentence level by Victor Chahuneau

Language Specific Projects / Manx

aspell-gv1about 14 years agoManx Gaelic dictionary for aspell
gaelg3about 2 years agoNLP resources for Manx Gaelic, mainly in support of the gv2ga MT engine

Language Specific Projects / Migmaq

migmaq-lessons1over 11 years agoRepository for website building Mi'gmaq language lessons

Language Specific Projects / Minderico

fredericajordarzambarino0about 12 years agoA web based game for mobile devices in minderico based in the "Who Wants to be a Millionaire" TV show

Language Specific Projects / Nishnaabe

Ojibway-iphone-app0about 11 years agoAn iPhone app with audio and images for learning the Ojibway language
OjibwayMap1about 11 years agoAn iPhone app with audio and images for learning Ojibway language and culture
nishanimate1about 11 years agoA desktop app to facilitate Nishnaabe-language acquisition via animations produced by the natural language processing of audio-accompanied text

Language Specific Projects / Oromo

hornmorpho5over 11 years agomorphological analysis and generation of amharic and oromo verbs and nouns. and tigrinya verbs

Language Specific Projects / Quechua

AntiMorfo5over 11 years agomorphological analysis and generation of Quechua nouns, adjectives, and verbs and Spanish verbs
Morphology, spellcheckerXFST and FOMA, plus OpenOffice plugin

Language Specific Projects / Sami

divvun-webdemo2about 3 years agosimple webdemo for divvun grammar checker.
GiellateknoA host of Sámi tools
Oahpa!A learning portal for Saami languages. Includes WordPress based, media rich lesson-based learning, and morphological and syntactic exercizes generated from the morphological and syntactic tools
NeahttadigisánitA morphologically sensitive dictionary, with modes for 'social media input' (which allows users to type a 'relaxed' version of the orthography ( will be recognized also as ), and also includes a JavaScript bookmarklet to offer click-to-read dictionary lookup functionality. Also available for . Giellatekno does a lot for other minority Uralic languages. Following are some keywords for CTRL+F friendliness:

Language Specific Projects / Scottish Gaelic

aspell-gd1about 14 years agoScottish Gaelic dictionary for aspell
briathrachan2almost 10 years agoThis is the source code to Briathrachan, a Gaelic-English dictionary app for iOS
gaidhlig3about 3 years agoNLP resources for Scottish Gaelic, mainly in support of gd2ga/ga2gd MT engines
gd-fcfg3over 14 years agoContext-free feature-based grammar of Scottish Gaelic in the NLTK format
gdbank4almost 2 years agoSome tools and resources for natural language processing of Scottish Gaelic.
hunspell-gd10over 3 years agoFiles for building Scottish Gaelic spell checkers

Language Specific Projects / Secwepemctsín

secwepemctsnem2almost 16 years agoA project to help people learn Secwepemctsín

Language Specific Projects / Somali

somorphSomali morphological and syntactic analyzers and generators built on XFST and VISL-CG Constraint Grammar. Up to date version checked in on repository
qaamuus.netmorphologically aware dictionary based on lexical resources found online, and the somali morphology

Language Specific Projects / Tigrinya

HornMorpho5over 11 years agomorphological analysis and generation of Amharic and Oromo verbs and nouns and Tigrinya verbs

Language Specific Projects / Uralic

UralicNLP71almost 2 years agoA Python library for processing Uralic languages (Finnish, Skolt Sami, Erzya, Moksha, Komi-Zyrian and so on). The library provides an easy programmatic access to Giellatekno resources such as FST morphology and CG disambiguators. Other functionalities include UD parser, API for the and interface to SemFi and SemUr semantic databases. The library is under active development and new features are added from time to time

Language Specific Projects / Zulu

UkwabelanaAn open-source morphological Zulu corpus

Backlinks from these awesome lists:

More related projects: