Building AI’s unified compute layer. We are hiring → modular.com/careers 🚀

Filter
Exclude
Time range
-
Minimum likes
Monday's community meeting showcases several standout community projects, including: • floki: an HTTP client for Mojo that works like Python's requests library, supported by the Modular Community Grant Program • noeira: an end-to-end physical AI stack in Mojo - physics, learning, perception and deployment, from simulator to robot, supported by the Modular Community Grant Program • warp: a look at building an async runtime to manage GPU synchronization calls across coroutines in Mojo Join us via Zoom at 10 AM PT: luma.com/sep-modular
1
15
1,662
This week in Chicago, Mojo found The Bean, ate a pretzel half its size, and got to hang out with the whole Modular team at our off-site. Want to come to the next one? We're hiring across Engineering, Product Management, Customer Engineering, and Developer Relations: modular.com/company/careers
1
2
29
1,914
.@clattner_llvm takes the #SnapdragonSummit stage for the first time as EVP Advanced AI Software @Qualcomm following the @Modular acquisition. Introducing the audience to the stack and it's role in the broader ecosystem.
3
17
1,003
GPU architecture | LLM Inference Handbook handbook.modular.com/kernel-… Add this to your LLM learning resource bundle. "Before writing or tuning GPU kernels, you need a working model of how a GPU runs code. Without it, suggestions like “increase occupancy” or "reduce shared memory bank conflicts" are just a set of rules to memorize. You don't fully understand when they apply and when they don't. This section explains modern GPU architecture at the level needed for kernel work. The details lean toward NVIDIA hardware because CUDA dominates much of the LLM inference ecosystem today. However, the core concepts apply broadly to AMD GPUs and other parallel accelerators as well."
4
99
616
29,777
Replying to @vivekgalatage
Thank you for the shoutout!
4
277
Optimizing large scale inference systems is what we do, so we decided to write down what we know. Our LLM Inference Handbook is a free reference covering TTFT, TPOT, goodput, continuous batching, chunked prefill, prefix caching, KV cache math, prefill-decode disaggregation, quantization, and more. It includes 20+ interactive visualizations, is updated continuously, and is open to PRs. handbook.modular.com
2
47
287
11,810
Modular retweeted
There’s no single answer for where AI goes next. At @modular’s ModCon, GV’s @davemuni joined leaders across AI and venture to talk about the next wave of opportunity. Dave’s view: the infrastructure is taking shape. Now, the exciting part is what founders build on top of it.
2
2
18
6,934
Wanted to do deep dive into inference engineering any good resource recommendations ?
10
137
1,057
65,819
Looking to get started with MAX? At ModCon 2026, Ehsan M. Kermani @ehsanmok and Bingfeng Xia went layer by layer through the stack: MAX Serve, the framework, and the Mojo kernel library, ending with a live agent bringing up a new model end to end. Start here: piped.video/watch?v=3H8Orjaa…
1
2
11
2,086
On the MAX side, audio joins text, vision, and image generation. The new audio_generation pipeline launches with MiniMax-Music3, producing 44.1 kHz stereo music from a style prompt and lyrics. MAX performance improved, too: up to 4.8x faster Gemma 4 decode attention and 7.9x faster MoE routing on NVIDIA B200, and up to 6.6x faster decode attention projections on AMD MI355. MAX changelog: max.modular.com/releases/v26…
1
5
615
We open sourced the Mojo compiler under Apache 2.0 at ModCon last month. Opening it to contributions was the top request we heard afterward, and it took some time to get the infrastructure right. It's ready now. We also migrated our internal issues to public GitHub issues, so you can read what the compiler team is working on this week. Mojo changelog: mojolang.org/releases/v1.1.0…
1
8
554
At ModCon this year, we asked five investors where the next wave of AI infrastructure capital is going: training, inference, or silicon? Five different answers, and one panelist said it's the wrong question to be asking. The discussion also covered whether chipmakers absorbing AI software means the infrastructure is maturing or getting ahead of itself, and what open weight models do to the valuation of a compute-heavy startup. Michelle Gonzalez (@M12vc), Liz Stein (@USITfund), Sam Fort (@dfjgrowth), Quentin Clark (@generalcatalyst) and Dave Munichiello (@GVteam) each closed with what they think founders should build right now. Full recording: piped.video/watch?v=5igY763x…
1
5
15
2,146
Fragmentation creates a tax across the entire AI ecosystem. At @PyTorch Conference this year, @clattner_llvm presents an alternative: one software stack to unite heterogeneous hardware. Join us in San Jose on October 20-21: hubs.la/Q04v4SL60 #PyTorchCon
New chips are shipping faster than ever before, but the ecosystem is being held back by having to rewrite and then re-debug the same code over & over again. @clattner_llvm CEO and Co-founder at @Modular and EVP of Advanced AI Software & Platforms at @Qualcomm, will deliver a keynote at PyTorch Conference North America about an open software platform for heterogeneous compute powered by Mojo and MAX. PyTorch has always been the place where the best models come together, and now there's a way to get those models onto all kinds of hardware. If you're interested in Al and compute, join us at the PyTorch Conference in San Jose, CA. Register now: hubs.la/Q04v4SL60 #PyTorchCon
1
5
63
6,365
Headed to Santa Clara today for #AIInfraSummit? Don’t miss @alisterburt's talk at 10:30 AM PT in Expo Theater 2: "Modular: Open Source, Open Cloud, Open Silicon." Stop by the Qualcomm booth (#206) anytime this week during expo hall hours to chat with the Modular team and catch a Modular Cloud demo.
3
15
1,626
If you looked at Mojo in 2024, liked it, and decided to check back later, later is now. Last month, the language hit 1.0, beginning a new epoch of stability for Mojo. The full Mojo language also went open source under Apache 2.0, including the compiler and more. Now is a great time to start building with Mojo. Clone the repo, build the compiler yourself, and learn from its unique MLIR internals and full commit history. At ModCon 2026, Brad Larson (Staff PM, Mojo) and Denis Gurchenkov (Senior Director, Mojo Compiler Engineering) walked through why Modular built a language at all, what stability means in the 1.x series, and a hint of what's next: piped.video/watch?v=a7YdrWOf…
5
11
76
6,102
We partnered with Hippocratic AI to integrate MAX into their inference pipelines running on NVIDIA B300 GPUs. Benchmarked against an existing SGLang deployment on 400B+ parameter models, MAX delivered sub-500ms mean time to first token and approximately 30% faster P99 end-to-end latency. Every millisecond matters in real-time voice, and these gains compound rapidly at the scale of Hippocratic AI's Polaris system.
1
4
655
.@hippocraticai's health agents call tens of thousands of patients a day. Each conversational turn has to finish in about 800ms or the call stops feeling human. That constraint is why they came to us.
2
18
2,714
Today's AI software is fragmented. Every accelerator has its own compiler, kernel libraries, runtime, and development path. Developers absorb the cost of this fragmentation. In his ModCon 2026 tech talk, Abdul Dakkak, Chief Scientist at Modular, presents our alternative: a unified compute layer, flexible enough to extend to new models, modalities, and hardware. Abdul walks through our recipe for bringing up new hardware, demonstrated with three examples: 1. AWS Trainium, via our own team. Gemma 4 31B end to end. 2. Google TPU v6e, through our partner @HTECgroup. Their team brought it up with no LLVM backend to lower to, no prior knowledge of Mojo or MAX internals, and minimal support from us. 3. d-Matrix Corsair, implemented by @dMatrix_AI's own team on their own stack. Repo access Tuesday, working matmul by the following Monday. Thanks to HTEC and d-Matrix for their collaboration, and to Mihailo for presenting HTEC’s learnings. Read about the HTEC collaboration: htec.com/insights/media-cove… Watch the full talk: piped.video/watch?v=hYjKbTAG…
1
5
22
1,641