Skip to content
#

text-normalization

Here are 87 public repositories matching this topic...

Text and orthography toolkit for 869 African languages: grapheme-to-phoneme (native spelling or IPA), conversion between any two languages' orthographies, one shared cross-language spelling, and a reversible plain-ASCII form. Non-Latin scripts romanised. Pure Python, no dependencies.

  • Updated Sep 29, 2026
  • Python

Text preprocessing and PII anonymisation for NLP/ML. ONNX NER ensemble, language detection, stopword removal. Built for statistical ML and language models.

  • Updated Sep 19, 2026
  • Python

A modern, actively maintained contractions library. Expands English contractions (you're → you are) with improved performance, type safety, and features like bulk dictionary imports and JSON loading. Includes 100% test coverage, full type hints, and works with both pip and uv.

  • Updated Oct 10, 2026
  • Python

Русская озвучка своим голосом: слой правки текста (ударения, бренды, твёрдая «э», аббревиатуры, транслит) и способ увидеть его работу ДО синтеза. CLI поверх F5-TTS. Russian TTS text layer + voice cloning CLI.

  • Updated Sep 19, 2026
  • Python

Add this topic to your repo

To associate your repository with the text-normalization topic, visit your repo's landing page and select "manage topics."

Learn more