Hongyi Liu retweeted
Flash-KMeans was only the beginning. Today, from the Flash-KMeans team, we are releasing FlashLib — a GPU library for fast, predictable, agent-ready classical ML operators. Up to 26× on KMeans, 19× on KNN, 40× on HDBSCAN, 208× on TruncatedSVD, 47× on PCA, 147× on exact t-SNE, and 49× on MultinomialNB over state-of-the-art (cuML). Blog: flashml-org.github.io/ Code: github.com/FlashML-org/flash…
47
228
1,600
873,152
Hongyi Liu retweeted
𝗞-𝗺𝗲𝗮𝗻𝘀 𝗶𝘀 𝘀𝗶𝗺𝗽𝗹𝗲. 𝗠𝗮𝗸𝗶𝗻𝗴 𝗶𝘁 𝗳𝗮𝘀𝘁 𝗼𝗻 𝗚𝗣𝗨𝘀 𝗶𝘀𝗻’𝘁. That’s why we built Flash-KMeans — an IO-aware implementation of exact k-means that rethinks the algorithm around modern GPU bottlenecks. By attacking the memory bottlenecks directly, Flash-KMeans achieves 30x speedup over cuML and 200x speedup over FAISS — with the same exact algorithm, just engineered for today’s hardware. At the million-scale, Flash-KMeans can complete a k-means iteration in milliseconds. A classic algorithm — redesigned for modern GPUs. Paper: arxiv.org/abs/2603.09229 Code: github.com/svg-project/flash…
36
199
1,755
317,254
Hongyi Liu retweeted
🚀 We just set a new SOTA for LLM inference acceleration with speculative decoding. By corralling a band of specialist drafters, we got 4.99× on Llama-3.1-8B-Instruct, 4.93× on Qwen-32B — beating EAGLE3 by nearly 2x. No gimmicks. Just careful math + solid engineering. 🧵1/
14
46
325
36,136
Hongyi Liu retweeted
🚀Excited to share our latest #EMNLP2024 work on benchmarking the long context ability with KV Cache compression across RNN-based architectures, token eviction, prompt compression, and quantization. We also provide an easy-to-use codebase (it also has my favorite WoW quote 😉). Feel free to give it a try and ⭐ it if you find it useful! 📄 Paper: arxiv.org/abs/2407.01527 💻 Code: github.com/henryzhongsc/long… Some interesting findings/suggestions include: 1️⃣ Maintaining an uncompressed prefill process is essential for performance, especially with harder tasks. 2️⃣ Combining RNN-based models with attention significantly enhances long-context capabilities. 3️⃣ In "needle-in-a-haystack" evaluation for recent LLMs like Llama-3, we should use longer needles (like 64 digits) since these models tokenize multiple digits into one token. More results and insights can be found in the paper! Kudos to all collaborators: @jiayiy, Hongyi Liu, @henryzhongsc, @YuNengChuang, Songchen Li, Guanchu Wang, Duy Le, @serendip410, Vipin Chaudhary, @ZhaozhuoX, @ziruirayliu, @huxia
1
10
41
10,249
So excited to share our recent work in @USENIXSecurity! It is an honor to be awarded distinguished paper and thanks to my team for the great work! 🎉If you are at the conference, please come to our session on Friday morning(session 2 track 6) and take a look😀 #usesec23
Congratulations to our USENIX Security '23 Distinguished Paper Award winners! bit.ly/usenixbestpapers #usesec23
7
716
Rice CS Ph.D. student Hongyi Liu, faculty member Ang Chen & team use a novel approach to deploy cloud devices as security enforcers. Liu will present their findings, Remote Direct Memory Introspection, at the 2023 @USENIXSecurity conference. bit.ly/3XqJlrE
1
6
678