وماذا عليك لو ذهبوا هُم بالدنيا، وأنتَ ذهبت بالقرآن؟ | AI Researcher @ KAUST - Kaggle Competitions Master (Highest Rank: 236th in the World)

🚨 New preprint GLIE: Generative Late-Interaction Embeddings for Visual Document Retrieval 🔵 TL;DR: Late-interaction retrievers store ~1,000 vectors per page. Turns out those ~1,000 vectors have only ~5 degrees of freedom, so a handful of stored vectors plus a tiny shared decoder can carry a page. So: store those projected vectors, use them as an index, and regenerate the full embedding set for the top candidates only, then rerank :)
3
13
92
64,805
Almost there...😁 1) (late-interaction): Generative Late-Interaction Embeddings 2) (spatial interpretablity): ⏳️ 3) (long context / retrieval / kv cache): ⏳️ 4) (long context / agentic): ⏳️ 5) (benchmarks saturation): ⏳️
I have a few research projects i have been working on with some geniuses in the past weeks including topics in: late interaction / interpretablity / retrieval / long context reading / benchmarks / RL post-training.... You are not ready for what i will post this month :)
7
1,032
Mohamed retweeted
#KostasThoughts: Seeing your former trainees succeed in life after the lab is incredibly rewarding.
1
2
45
3,405
Happy to say it got accepted to #NeurIPS2026 🎊🎊
🚨 New preprint: GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs 🔵TL;DR: new sub-quadratic training-free inference method for video VLMs. ➡️3.36× less compute, no accuracy loss, fully interpretable on the cell-level, adaptive frames selection using VLM reasoning rather than CLIP-like tricks! 🔥Spoiler alert: Surprisingly, question difficulty alone tells you how many frames a long-video VLM needs to answer it! We turn this into a closed-form rule!!
1
2
21
975
The vectors of late interaction models are highly redundant and it also has an exploitable geometry So why not learn a decoder to exploit it and regenerate the original vectors from a set of pooled ones? It allows to store only a small part of the vectors, use them to generate candidate and then decode to rerank! I think it starts to bridge the gap between ColBERT and MICE (cc @_VictorMorand_) and it is definitely a nice avenue!
🚨 New preprint GLIE: Generative Late-Interaction Embeddings for Visual Document Retrieval 🔵 TL;DR: Late-interaction retrievers store ~1,000 vectors per page. Turns out those ~1,000 vectors have only ~5 degrees of freedom, so a handful of stored vectors plus a tiny shared decoder can carry a page. So: store those projected vectors, use them as an index, and regenerate the full embedding set for the top candidates only, then rerank :)
2
6
72
4,823
Mohamed retweeted
Generative Late-Interaction Embeddings for Visual Document Retrieval Compresses multi-vector document embeddings to a few per page via a learned code that regenerates the full set for reranking, beating prior compression without retraining the encoder. 📝arxiv.org/abs/2609.11808
1
9
68
5,143
🚨 New preprint GLIE: Generative Late-Interaction Embeddings for Visual Document Retrieval 🔵 TL;DR: Late-interaction retrievers store ~1,000 vectors per page. Turns out those ~1,000 vectors have only ~5 degrees of freedom, so a handful of stored vectors plus a tiny shared decoder can carry a page. So: store those projected vectors, use them as an index, and regenerate the full embedding set for the top candidates only, then rerank :)
3
13
92
64,805
📄 Paper: arxiv.org/abs/2609.11808 💻 Code: github.com/mohammad2012191/G… Huge thanks to the great team: Talal Aloushan, Rose @Quantumleap___ , Jana Shatta, Mohammed Alhassan, Leen Alrehaili, and supervisors Dr. Tanveer Hussain & Prof. @ProfNaeemKhan. Special thanks to @KAUST_Academy & @Dr_S_Albarakati for their absolute incredible support🫡 and special thanks to @lateinteraction and @antoine_chaffin for their continuous feedback!
1
7
400
Mohamed retweeted
This is notable. DeepSeek, a lab usually first to pioneer novel algorithms and architectures, is saying that at this point, the ROI of improving data quality far exceeds that of working on novel post-training algorithms. I think this has already been true for some time for non-lab practitioners. If you're doing llm post-training, 80% of your effort should go into looking at your data. This means: - Hiring experts to dig through your RL tasks - Sifting through rollouts and sft data by hand to remove suspicious samples. Make sure all tasks are actually passable. - Making sure your data is diverse in both difficulty and category.
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
57
151
1,733
209,445
I have a few research projects i have been working on with some geniuses in the past weeks including topics in: late interaction / interpretablity / retrieval / long context reading / benchmarks / RL post-training.... You are not ready for what i will post this month :)
16
1,768
So K3, with one simple question: how do you even train a 3T parameters model? Even though the paper is not out, there is enough public information to get a relatively consistent picture. Just piling on experts won't make an economically viable model. What worked and made training and inference business profitable is highly sparse models with high-margin: initial investment in compute infrastructure is recouped through expert parallelism. Now as you get way beyond 1T parameters, you start running into physical limits of the available infrastructure. 896 experts won't fit on one node, even when only 16 are activated at generation time, you need to start splitting weights and latency starts to leak on all sides. First, K3 is a LatentMoE. The technique is not invented by Moonshot but by Nvidia (arxiv.org/pdf/2601.18089) and it's fundamentally a communication trick: instead of routing on the full hidden state, we down-projec, route, run the experts in that latent space, and up-project back. It's a fixed cost so that you can grow an indefinite amount of experts and still bound communication and bandwith. Now the critical part: Nvidia approach was focused on inference and Moonshot's approach is a *Stable* LatentMoE oriented toward training. Extreme sparsity means that each expert sees very few tokens per batch: this new stable implementation seems to fix the load-balancing estimator getting incredibly noisy as sparsity grows. Quantile balancing addresses a common MoE equilibrium issue that becomes explosive: some experts are too popular, too many token gets routed to them and then other experts just have to wait, idle and starving. K3's announcement explicitly said they anticipated expert paralllelism at inference time as a shaping training constraints. With current loading methods, we don't know how many tokens each expert got until they're routed, so buffers can't be sized ahead and you need a host round-trip to find out. Instead quantile loading maximizes total routing score subject to each expert taking exactly mk/n tokens. Roughly, each expert takes exactly its quota by constructio, then every buffer is known at compile time and the sync disappears. Attention Residuals seems to be fundamentally a depth trick already published by MoonShot but never validated at scale (the AttnRes paper only trains at 48B). Attention is pointed at layers instead of tokens. Normally every layer dumps its output onto a shared pile (the residual stream) and every later layer reads the whole pile back, everything added with equal weight. AttnRes lets each layer attend over the earlier layers and pull out what it actually wants. Same shape of deal as LatentMoE: a small fixed cost, under 5% overhead for about 25% of the gain. DeepSeek's mHC went in a similar direction by widening the residual stream until it is effectively doing attention across layers, except implicitly, through a chain of matrix products that can explode or collapse, which is why mHC needs sigmoids and doubly-stochastic normalization to hold it together. AttnRes doesn't fix that machinery, it avoids needing it: write the attention explicitly and softmax gives you the guarantees for free. We have much less information at this stage about the other new components (head muon, gated MLA, SiTU (though last one is likely needed for MXFP4?). Yet they all seem to proceed from the same philosophy shared by Su Jianlin in an April post (kexue.fm/archives/11729): "small-model-safe does not mean big-model-safe". Neither LatentMoE or Quantile Balancing made already much sense at smaller scale. They are becoming an integral part of the training and inference stack as we move up to an entirely different dimension of scaling and infrastructure deployment.
10
70
747
82,462
something weird i thought about today: what if an LLM could predict its own emotion as it outputs text? Sometimes a model says something like "LOOK AT THESE RESULTS" which obviously shows excitement, but how to measure such a thing numerically? I know I can just ask it about its emotion and it'll tell me, but that feels more like an external measure, and I care about something internal, because emotions feel like a property that comes bundled with the words, not something measured from the outside. we as humans have feelings and emotions while we talk, so what if we could inject something like that into the text itself? Imagine labels: angry, excited, ... represented as [0,1,1,...] ALONG each word. Then the model predicts that distribution after every word along with predicting the next word. I believe this might make huge unlocks in: AI writing, because the model uses both signals to decide what comes next. I feel it is just interesting to see how the emotion distribution changes as the text goes, particularly for @claudeai hahaha
1
7
909
if u wanna have a bit of fun :)
بدأت المنافسة 🚀 انضم إلى تحدي "مُرَقِّم" لترقيم النصوص العربية الآن! رابط الانضمام: forms.gle/DhD1UMuSLPW1HGd76 #تحدي_دال
9
3,610