A 200K-vocab tokenizer for 44 languages, built so non-Latin scripts don't pay the fertility penalty they take in standard BPE.
ProductionStandard BPE tokenizers are trained on Latin-heavy corpora and it shows: Hindi, Tamil, or Bengali text routinely needs 2-4x more tokens per character than English for the same content. That's slower inference and worse context utilization for exactly the languages we care most about.
Not every idea worked. Morphology suffix protection caused a Hindi fertility regression before being fixed back to neutral. Iterative reallocation turned out to be a no-op on small corpora because the vocabulary was already saturated. Both are documented, not smoothed over.
Head-to-head against our own Hyper-Token on FLORES-200, PU-Tok comes out ahead on Cyrillic, Greek, Armenian, Thai, Lao, Khmer, and Amharic — scripts with dedicated handling in its pipeline.
© 2026 Cybergeon Technologies. All numbers on this site are re-verified against real code before publishing.