Developer Advocate @NebiusAI. Building with open models | coding agents | LLM inference

San Jose, CA, USA
Filter
Exclude
Time range
-
Minimum likes
Replying to @hanrui_w
btw @hanrui_w - those clustered release dots toward the end of the timeline? That acceleration is all thanks to your team 👏
1
1
28
an animation of 2026 model releases in @nebiustf 👇 I love creating visualizations. And now that the tools and models have gotten so good, creating something like this is surprisingly accessible. :-) (will post prompt + assets soon)
1
2
13
2,309
i love when the models go UP (intelligence) and to the LEFT (lower pricing) viz : sujee.github.io/practical-ll…
Qwen3.8-27B is now live on Nebius Token Factory. A compact 27B dense model for coding, research, and agent workflows, with a focus on planning and completing tasks across multiple steps. Start building: tokenfactory.nebius.com/endp…
2
184
Replying to @Astrodevil_
very nice 👏 Qwen3.8-27B is indeed a powerful little model!
1
1
24
Replying to @Astrodevil_
Very cool... Did you give it any pointers like GitHub repo etc?
1
1
27
Replying to @irastech
hehe, i started this as a fun side project.. but learned a lot about LLM behavior while doing this :-)
1
6
and @demian_ai when are you forming the quant fund? ;-)
1
534
$10,471.69 times 1000x worth idea right here 👏
1
1
928
Try it for yourself 🐍 sujee.github.io/llm-snake-ar… Go burn some tokens :-) Code: github.com/sujee/llm-snake-a…
1
33
Friday funsies 🐍 LLM Snake Arena: GPT-6-Astra vs Claude Opus 5.5 Does it tell us something about the models? Maybe. Fun to watch? Most definitely 😀
2
6
193
for me this is the best part... not another skill to install and learn. just a straightforward prompt.. and you get an amazing video.. did you have to iterate over to get the final video @nutlope ?
34
Replying to @nutlope
> I love open models and use them a lot, but I'm also a big believer in using the right model for the job 👆 Well said! Lately I've been using Claude Code to edit videos, and it is really, really good. As a fan of open models, I keep coming back to the same conclusion: use the best tool for the job.
1
1
111
@deepseek_ai V4.1 Flash is now available on @nebiustf Here’s a quick demo in the Playground, including a speed test and a vision example. Links and credits in the reply 👇
2
1
5
116
haha, not sure who is more impressive.. human @DhruvDiddi doing pushups :💪 or robot 🤖 with the moves!
1
2
45
Want to try open models for coding? Keep your coding agent - Claude Code, Codex, OpenCode, Cline - and just swap the model underneath. In this short talk: * open vs closed models * why the coding harness matters * what coding tasks cost * quick setup demo Watch: piped.video/watch?v=gjNcSo1M…
3
8
278
2 really good open models - GLM-5.3 and DeepSeek-V4-Pro-0813 now on @nebiustf pushing intelligence while being pretty economical! viz from Token Factory model visualizer : sujee.github.io/practical-ll…
Two new models for coding and agent workflows are now live on Nebius Token Factory. DeepSeek V4-Pro-0813 is the official V4 Pro release, built for coding agents that use tools, reason through complex problems, and work across multiple steps. GLM-5.3 brings Z. ai’s latest post-training improvements for complex software engineering, from planning changes across a repository to carrying long-running agent tasks through to completion. Try both through an OpenAI-compatible API, with Token Factory handling the serving infrastructure. Bring your own tasks and see which fits your workflow. Start building: tokenfactory.nebius.com/mode…
2
6
373
Had a fun time geeking out and building with @dN0t and @devopsjacquie - using open models for coding -- they are surprisingly good (and cheaper) - using Hermes to find vintage auto parts! - new models in @nebiustf Get started by singing up to the Builder Program : lnkd.in/gkUXHpMS And share with us what you build. nitter.net/i/broadcasts/1rGmqpNjX…
1
6
153
Went to buy a couple of 4TB portable SSDs and almost fell off my chair looking at the prices 😳 Damn. See screenshots 👇 @demian_ai has been writing about the memory/storage crunch driven by the AI boom. It is finally sinking in for me :-) A few of his posts: - nitter.net/demian_ai/status/20644… - nitter.net/demian_ai/status/20092… - nitter.net/demian_ai/status/20766…
$PENG turned scarcity into revenue. Samsung made memory stretch. Micron locked down wafers. Korea broke the wrapper. this week, the book looked flat but the stack was being repriced ↓ Under a flat headline, the market made a clear choice: pay for what can turn scarcity into revenue now. Storage, networking, AI cloud and cooling led. HBM / packaging and power / grid lagged. $PENG, Samsung, Micron and Korea explain the rotation. REVENUE NOW - $PENG Penguin was the receipt. Q3 revenue grew 48%. Integrated Memory more than doubled. Operating income rose more than 5x. The point is not just that memory is tight. Penguin gets paid to turn memory, clusters and infrastructure software into a working AI system. Scarcity matters but making it usable pays sooner. MEMORY ELASTICITY - SAMSUNG Samsung's Blackwell test added a 1TB CXL pool for KV-cache offload. When local DRAM filled, throughput collapsed. The CXL-backed system kept running near DRAM speed (approximately 92% of DRAM performance in multi-GPU configurations). Not because CXL replaces HBM but because not every byte needs the fastest, most expensive tier. The hierarchy is widening: HBM → DRAM → pooled CXL → storage The caveat: this required host changes, a custom in-house kernel and modifications to the LMCache stack. Engineering proof, not plug-and-play adoption. But it shows where the next memory trade may form: deciding where each byte belongs. SUPPLY SECURITY - MICRON Micron put $500M behind GlobalWafers' 300mm Texas plant and signed a 10-year supply agreement. That is more revealing than another demand forecast. When the customer finances the supplier, the bottleneck has moved from the presentation deck to capital allocation. AI memory profits are being recycled into the layer beneath memory: wafers, materials and geographic resilience. RACK TIME VS GRID TIME Transformer queues now stretch beyond three years. High-voltage breakers are not far behind. Yet cooling rallied while Power & Grid fell. The physical constraint did not vanish. The market separated two clocks: Cooling gets installed with the rack. Grid equipment gets paid after permitting, financing and construction. Same density problem. Very different route to cash. DEPLOYABILITY BEAT DESTINY Networking and retimers outperformed photonics / CPO. $ANET led; $CRDO gained. That is not a verdict against optics. CPO may still be the architectural destination. But copper, retimers and systems already shipping can earn during the transition. The market paid for what can deploy before the perfect end state arrives. SAME ASSET, DIFFERENT PLUMBING SK hynix's U.S. ADR jumped on debut. The next trading day, its Seoul shares fell more than 15%, Samsung fell sharply, the KOSPI fell ~9% and trading halted. Reuters calculated a roughly 37% ADR premium after the rout. Same company, same HBM exposure, but different access, liquidity, leverage and flows. Korea did not prove that the HBM thesis was broken. It proved that the wrapper can overpower the asset. THE THREE CLOCKS 1. Revenue now: storage, networking, AI operations, cooling. 2. Scarcity later: power, fabs, packaging, substrates. 3. Elasticity: CXL, retimers, liquid cooling, orchestration. That is the map for this week: $PENG = turn scarcity into revenue. Samsung = make scarce memory go further. $MU = secure the layer underneath. Korea = price the instrument, not only the asset. The bottleneck did not disappear. The market started asking a harder question: How long until it becomes cash? Full weekly: aibottlenecks.app/alpha/week…
1
2
8
3,136