happy to announce that we've gotten rid of tokenizers!
especially excited with what we've replaced them with: end-to-end trainable modules that not only learn to group characters into (sub)words, but can iterate to group words into phrases and further higher-order concepts
see @sukjun_hwang's thread for more details 👇
Tokenization has been the final barrier to truly end-to-end language models.
We developed the H-Net: a hierarchical network that replaces tokenization with a dynamic chunking process directly inside the model, automatically discovering and operating over meaningful units of data
Jul 11, 2025 · 4:09 PM UTC
13
50
758
82,286
on a more personal note, this was my first real taste of ml research, and a super invigorating project at that. i had so much fun and learned an incredible amount from working with the amazing @sukjun_hwang. (and @_albertgu is pretty great too!)
2
48
2,352











