happy to announce that we've gotten rid of tokenizers! especially excited with what we've replaced them with: end-to-end trainable modules that not only learn to group characters into (sub)words, but can iterate to group words into phrases and further higher-order concepts see @sukjun_hwang's thread for more details 👇
Tokenization has been the final barrier to truly end-to-end language models. We developed the H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data

Jul 11, 2025 · 4:09 PM UTC

13
50
758
82,286
on a more personal note, this was my first real taste of ml research, and a super invigorating project at that. i had so much fun and learned an incredible amount from working with the amazing @sukjun_hwang. (and @_albertgu is pretty great too!)
2
48
2,352
Sort replies: Relevant Recent Liked
Replying to @fluorane
Hello, this is an interesting paper. I hope in the future there will be work in this area that investigates different choices of text encoding than UTF-8, or the use of different encodings for the inputs and outputs.
1
4
1,237
Replying to @fluorane
INSANE work for the team!!
3
1,073
Replying to @fluorane
incredible work! 🔥
1
415
Replying to @fluorane
holy shi
1
407
Replying to @fluorane
this is awesome!
1
646
Replying to @fluorane
congrats!!
1
534
Replying to @fluorane
Tokesux!
5
923
Replying to @fluorane
wow soooo cool!!!
2
389
Replying to @fluorane
Promising work
1
586
Replying to @fluorane
Replacing tokenization with a learned hierarchical chunking process is interesting because it asks the model to discover the units it needs instead of inheriting them from a fixed preprocessing step.
1