yalign

Sentence aligner

Automates the process of extracting parallel sentences from comparable corpora to aid in statistical machine translation

A sentence aligner for comparable corpora

GitHub

127 stars
16 watching
31 forks
Language: Python
last commit: over 10 years ago
Linked from 1 awesome list


Backlinks from these awesome lists:

Related projects:

RepositoryDescriptionStars
lowresourcelanguages/champollionA toolkit providing ready-to-use parallel text sentence alignment tools for multiple language pairs.18
braunefe/gargantuaSoftware tool to align sentences across multiple languages using unsupervised alignment methods12
jcgood/rosetta-panglossA Python library that uses machine learning and natural language processing to improve translation accuracy by aligning source and target languages0
lowerquality/gentleA tool for aligning speech with text by analyzing audio and providing an output transcript1,471
lewisgaul/zig-nestedtextA parser library for a human-readable data format based on YAML13
prosodylab/prosodylab.alignertoolsA package of scripts to prepare data for alignment in speech processing12
montrealcorpustools/montreal-forced-alignerA command-line utility for aligning audio data with written text based on pronunciation rules.1,364
tanloong/interlaced.nvimA plugin for aligning bilingual parallel texts by re-positioning text and applying highlighting.7
prosodylab/prosodylab-alignerTools for aligning laboratory speech production data to forced audio alignment using HTK and SoX.333
talschuster/crosslingualcontextualembEnables alignment of word embeddings across multiple languages to facilitate cross-lingual text analysis and machine learning tasks99
thom1729/yaml-macrosA macro system for YAML files powered by Python21
zhoux85/stalignerTool for aligning and integrating spatially resolved transcriptomics data using machine learning algorithms29
kaaaaaaaaaaai/paragraph-with-alignmentProvides a paragraph tool with alignment options for the Editor.js text editor framework.44
cmesher/inuktitutalignerdataScripts for aligning laboratory speech production data in Inuktitut3
vchahun/gv-crawlAutomates text extraction and alignment from Global Voices articles to create parallel corpora for low-resource languages.9