💫 Industrial-strength Natural Language Processing (NLP) in Python
-
Updated
Sep 30, 2026 - Python
💫 Industrial-strength Natural Language Processing (NLP) in Python
Easy token price estimates for 400+ LLMs. TokenOps.
All the slides, accompanying code and exercises all stored in this repo. 🎈
👑 spaCy building blocks and visualizers for Streamlit apps
Trankit is a Light-Weight Transformer-based Python Toolkit for Multilingual Natural Language Processing
Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).
The official code 👩💻 for - TOTEM: TOkenized Time Series EMbeddings for General Time Series Analysis
Toonify: Compact data format reducing LLM token usage by 30-60%
[NeurIPS 2024]OmniTokenizer: one model and one weight for image-video joint tokenization.
Rule-based token, sentence segmentation for Russian language
[Paper][AAAI 2025] (MyGO)Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity Representation
Simple multilingual lemmatizer for Python, especially useful for speed and efficiency
[CVPR '26] SceneTok: A Compressed, Diffusable Token Space for 3D Scenes
Code for the paper "Fishing for Magikarp"
Fast bare-bones BPE for modern tokenizer training
Implementation of the LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens Paper
A unified tokenization tool for Images, Chinese and English.
Code for Zero-Shot Tokenizer Transfer
Single-stage End-to-End Training for Tokenization and Generation
Implementation of the GBST block from the Charformer paper, in Pytorch
To associate your repository with the tokenization topic, visit your repo's landing page and select "manage topics."