Research Scientist working on Deep Learning, Time Series Forecasting, Reinforcement Learning and HPC.

Berlin, Germany
Kashif Rasul retweeted
⚡ Working on a major speedup for AsyncGRPO in TRL: training steps are up to 3.5× faster so far! AsyncGRPO already has several packing optimizations. One packs samples together and balances the attention cost across DP ranks. You start with the same prompt G times. In multi-turn rollouts on a harness, the conversation can also start to fork: most of the history remains identical, while a tool result or a new turn changes the suffix. That gets expensive with GRPO because rollouts can share many long prefixes. You start from the same prompt G times, and multi-turn rollouts on a Harness can also share most of their conversation history (thanks to message-level tokenization). Each rollout is treated as its own sequence and packed on a single row, subject to a predefined token budget per DP-rank. Those rows are forwarded independently, even when most of their tokens are identical. Tree packing instead turns the rows into a prefix tree. Shared prefixes are stored and forwarded once, while we use FlexAttention to ensure that every token still sees exactly the context it would have seen in its original rollout. The loss is still defined on the original rollouts, so every trained token keeps its own target, advantage, and old log-prob. The video shows the algorithm pretty well 👇
5
9
84
6,819
Kashif Rasul retweeted
Happy to officially bring you the new SOTA `tokenization` library. We focused on all languages, multi-thread scaling, minimal package size and memory usage. huggingface-tokenizers-v1.st… We're thankful for the players of this ecosystem that have pushed us to give the best we could!
6
39
298
20,960
Kashif Rasul retweeted
Happy to announce the newest 🤗 PEFT release, v0.21.0 🎉 github.com/huggingface/peft/… Check the thread below for a few highlights of this release.
1
3
9
635
Kashif Rasul retweeted
Can you predict next week's weather yourself? We just wrote a blog post that explains how! ⛅️🌍🌧️ Running weather forecasting required having a supercomputer and CS/physics Phd. Now, AI-based models need fewer resources and are available to anyone on @huggingface. So, prediction the weather should be accessible to all. But only a few labs can actually run these models because there is no standardized documentation guiding you from the model weights to an actual forecast. Together with @EarthmoverHQ, we wrote a blog post and a demo walking you through that: huggingface.co/blog/hugging-… The blog covers the code for fetching initial conditions from the Earthmover Marketplace, building a model's input batch, running inference and comparing the forecast against historical climate data. You can also run Aurora (Microsoft), WeatherNext2 (Google) and AIFS (ECMWF) on our demo space. This is a first step to make AI-based weather forecasting models more accessible, and we'd love to build more of that. Tell us what's missing or your feedback in the blog comments!
8
25
124
19,744
Kashif Rasul retweeted
Nvidia NeMo AutoModel now supports DFlash2 training 🚀 DFlash 2 keeps the best part of block-parallel drafting while improving candidate selection and refinement, delivering 25–37% higher acceptance length and 27–39% higher single-request throughput than DSpark on the published Qwen3.8-27B benchmarks. AutoModel now includes the DFlash 2 model, trainer, recipe, Qwen3/Qwen3.5 support, and multi-GPU training, so you can train your own drafter instead of relying on a prebuilt checkpoint. github.com/NVIDIA-NeMo/Autom… Many thanks to @krasul and @Chonghan_Liu for adding the support!
4
29
2,962
Kashif Rasul retweeted
Very cool open source model available on Hugging Face: Samudra 2 runs multi-year ocean forecasts (temperature, salinity, sea surface height,...) at resolutions down to 1/4° and on a single GPU. If you're working on similar models, feel free to reach out. We're thinking about build a benchmark for ocean and atmospheric forecast models 🌍🌊 Try it here: huggingface.co/spaces/huggin… Model weights: huggingface.co/M2LInES/Samud…
3
4
24
1,132
Kashif Rasul retweeted
Same reward curve. A quarter of the clock! 🚅 TRL's AsyncGRPOTrainer decouples rollouts from training: a background worker streams completions from a vLLM server while the training loop consumes them. The bench is the well-crafted sail/Sanity-Test-R1D-1.5B. 3.8× on the clock. 1.9× on GPU-hours.
1
3
17
2,390
Kashif Rasul retweeted
you can now train DiffusionGemma (a block-diffusion LLM) in TRL! and we're sharing an example for it TRL trainers are made to be easily extended and adapted to different real-world use cases. in this one, with a single method overridden in SFTTrainer (compute_loss), you can train this model example below
1
2
11
702
Kashif Rasul retweeted
I built a small ECHO demo on TRL + OpenEnv❗️ the nice part: TRL's trainers are easy to adapt, so a small trainer subclass does the job ran it locally and the world-model signal shows up: env-prediction improves with ECHO (pink), flat with GRPO (teal)
ECHO (Environment Cross-entropy Hybrid Objective) demo support just landed in OpenEnv, and it's a cool idea: train agents to learn a world model almost for free original paper by @VaishShrivas, Piero Kauffman, Ahmed Awadallah and @DimitrisPapail @MSFTResearch when an agent acts in an env, a rollout has 2 sides: what the agent writes and what the env writes back
2
6
27
4,246
gave an llm access to a thermal camera pointing to the raspberry pi its running from and it started to do experiments on turning the fan on and off to check for temp changes... 🤯
2
2
36
8,906
Here the LLM is turning on and off different GPIO pins to figure out where the fan is attached to and if its working:
2
10
694
It also ran different denoising filters to get better images:
167
This experiment gives new meaning to llm observability, where we can imagine the large llm can see the state of its datacenter or even the patch of satalite image of the datacenter and can perhaps take actions against env. impact? Food for thought. 🤔
2
1
6
537
Kashif Rasul retweeted
🔥 AITER's FA3, RoPE, and MegaBlocks kernels are now on the 🤗 Hub and integrated into Transformers. No model code changes required. Just enable Hub kernels and go. 🚀 Up to 1.63× faster GPT-OSS-20B inference on MI300. #AMD #ROCm #HuggingFace
3
2
13
1,620
Kashif Rasul retweeted
When it comes to fine-tuning, LoRA is dominating. But is it the best option? We investigate this in our latest blog post on HF. You'll learn which PEFT techniques work best on our benchmarks and how to choose the PEFT technique that is best for your use case.
1
2
34
9,458
Kashif Rasul retweeted
Deep dive into FNS: building a tokenizer that chunks text efficiently but has character level resolution! FNS augments the loss with character level signal at training time while at inference time you can decode single characters. Deep dive here: huggingface.co/spaces/Huggin…
6
11
45
4,613
Kashif Rasul retweeted
OpenEnv already ships 🚢 with a ready-to-deploy RLM environment on free HF Spaces Drop "Attention Is All You Need", write code that spawns parallel LLM calls → ✅ answer in 4.2s Run GRPO (TRL) → model learns to write that search strategy itself 👀@lateinteraction @a1zhang
2
10
47
4,491
Kashif Rasul retweeted
Earlier this month, Apple introduced Simple Self-Distillation: a fine-tuning method that improves models on coding tasks just by sampling from the model and training on its own outputs with plain cross-entropy and… it's already supported in TRL, built by @krasul. you can really feel the pace of development in the team 🐎 paper by @onloglogn, @richard_baihe, @UnderGroundJeg, Navdeep Jaitly, @trebolloc, @YizheZhangNLP at Apple 🍎 how it works: the model generates completions at a training-time temperature (T_train) with top_k/top_p truncation, then fine-tunes on them with plain cross-entropy. no labels or verifier needed you can try it right away with this ready-to-run example (Qwen3-4B on rStar-Coder): github.com/huggingface/trl/b… or benchmark a checkpoint with the eval script: github.com/huggingface/trl/b… one neat insight from the paper: T_train and T_eval compose into an effective T_eff = T_train × T_eval, so a broad band of configs works well. even very noisy samples still help want to dig deeper? paper: huggingface.co/papers/2604.0… trainer docs: huggingface.co/docs/trl/main…
5
36
224
37,889
Kashif Rasul retweeted
Today, we released PEFT v0.19.0 and it's a big one. Not only did we add 9 new PEFT methods, the release also contains a bunch of improvements to make PEFT more useful. Check the thread for details:
2
2
13
598