Mehedi Ahamed retweeted
Excited to share that mm-ctx is now live on @huggingface Spaces! Try it in the browser via an interactive terminal without installing anything: vlm-run-mm-ctx.hf.space mm-ctx – fast, multimodal context for agents. LLM-based agents handle text fine, but as soon as a directory contains images, videos, or PDFs with visual content, they struggle to understand the full context. mm-ctx is meant to feel familiar: the Unix tools we already love (find/cat/grep/wc), rebuilt for file types LLMs can't read natively and designed to work with agents via the CLI. - mm grep "invoice #1234" ~/Downloads searches across PDFs and returns line-numbered matches - mm cat <document>.pdf returns a metadata description of the file - mm cat <photo>.jpg returns a caption of the photo - mm cat <video>.mp4 returns a caption of the video A few things we obsessed over: ⚡ Speed: Rust core for the hot paths 🏠 Local-first, BYO model: Uses any OpenAI-compatible endpoint: Ollama, vLLM/SGLang, LMStudio with any multimodal LLM (Gemma4, Qwen3.5, GLM-4.6V). 🔗 Composable: stdin + structured outputs 🤖 Drops into any agent via mm-cli-skills: Claude Code, Codex, Gemini CLI, OpenClaw. We’d love to hear your feedback! Especially on the CLI and what file types and workflows you would like to see next.
4
7
14
1,239
New Publication Alert Accepted at WACV'26- From Lightweight CNNs to SpikeNets: Benchmarking Accuracy-Energy Tradeoffs with Pruned Spiking SqueezeNet
From Lightweight CNNs to SpikeNets: Benchmarking Accuracy-Energy Tradeoffs with Pruned Spiking SqueezeNet Radib Bin Kabir, Tawsif Tashwar Dipto, Mehedi Ahamed, Sabbir Ahmed, Md Hasanul Kabir arxiv.org/abs/2602.09717 [𝚌𝚜.𝙲𝚅 𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙴𝚃 𝚌𝚜.𝙽𝙴]
4
Virtual Try-On is getting scary good... I think I just found my new favorite tool 📷 I was experimenting with VLM Run today to see if I could virtually "try on" a dress from a product photo onto an existing picture. Is this the future of online shopping?
1
1
7
AI can produce cute videos, commercials...Damn!!!
7
Amazing to see how (chat.vlm.run) is transforming advertisement industry.
7
Bring image to life using (chat.vlm.run). It doesn't just "see" video, it understands texture, movement, and minute details at a granular level.
6
Been testing tons of CV tools (Roboflow, SuperAnnotate, HF Spaces, etc.), but VLM Run (vlm.run) genuinely surprised me. Their new chat feature turned a simple factory blueprint into a full, realistic manufacturing workflow. Vision-driven execution at its best.
6