Developer Advocate @NebiusAI. Building with open models | coding agents | LLM inference

San Jose, CA, USA
Here’s a quick explainer on fine-tuning, distillation, and quantization - using LEGO® bricks 🧱 🟢 Fine-tuning: specialize a pretrained model with new data. 🟠 Distillation: train a student model to learn from a teacher. 🟡 Quantization: use lower-precision weights to reduce memory and bandwidth use. Video below 👇 More detail: sujee.dev/post/fine-tuning-d… What should I explain next with LEGO: LoRA, pruning, or speculative decoding? Drop your vote in the comments 👇
1
1
6
219
Friday funsies 🐍 LLM Snake Arena: GPT-6-Astra vs Claude Opus 5.5 Does it tell us something about the models? Maybe. Fun to watch? Most definitely 😀
2
6
181
Try it for yourself 🐍 sujee.github.io/llm-snake-ar… Go burn some tokens :-) Code: github.com/sujee/llm-snake-a…
1
25
Sujee Maniyam retweeted
I was testing DeepSeek V4.1 Flash on @nebiustf , and honestly, the model is pretty damn good. So I built a tiny real-time webcam app that takes multiple frames from the camera and uses the model to summarize what it sees. The whole thing was surprisingly smooth, and the results were quite impressive for such a small setup. Here’s a quick demo of it:
DeepSeek-V4.1-Flash is now live on Nebius Token Factory The model comes with native image understanding plus an architecture built for faster inference and higher throughput on long, input-heavy workflows. It combines native vision and a 1M token context with a 552B MoE that activates just 8B parameters for input and 16B for output. Try it out: tokenfactory.nebius.com/?mod…
2
4
13
1,985
2 really good open models - GLM-5.3 and DeepSeek-V4-Pro-0813 now on @nebiustf pushing intelligence while being pretty economical! viz from Token Factory model visualizer : sujee.github.io/practical-ll…
Two new models for coding and agent workflows are now live on Nebius Token Factory. DeepSeek V4-Pro-0813 is the official V4 Pro release, built for coding agents that use tools, reason through complex problems, and work across multiple steps. GLM-5.3 brings Z. ai’s latest post-training improvements for complex software engineering, from planning changes across a repository to carrying long-running agent tasks through to completion. Try both through an OpenAI-compatible API, with Token Factory handling the serving infrastructure. Bring your own tasks and see which fits your workflow. Start building: tokenfactory.nebius.com/mode…
2
6
371
As I was updating my model indexes, I noticed that @ArtificialAnlys has updated its Intelligence Index. The new AA Intelligence Index v4.3 introduces several significant changes, including a tougher set of evaluations. You can read about the changes here: nitter.net/ArtificialAnlys/status… One immediate effect: scores are lower across the board. That doesn't necessarily mean the models got worse. The benchmark suite got harder, so the scale has effectively shifted. At the top of the latest leaderboard: - Claude Fable 5.1: 53 = 66 → 53 = ▼ 13 - GPT-6 Astra: 53 = 61 → 53 = ▼ 8 And the top 3 open-weight models: 1. GLM-5.3: 45 2. Kimi K3: 44 (60 → 44 = ▼ 16) 3. GLM-5.3-Flash: 42 (57 → 42 = ▼ 15 ) The score changes for open models are pretty substantial. See the images below 👇 for the changes and the latest leaderboard. As they say, it's all relative 😀 What matters most is how models compare against each other under the same evaluation suite.
Announcing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5 Changelog (Index v4.2 → Index v4.3): ➤ Terminal-Bench: 2.1 → 4.0, completing our upgrade to the latest version of Terminal-Bench ➤ Replacing 𝜏³-Banking with AutomationBench-AA, our implementation of Zapier's business workflow automation benchmark We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5. Each change in v4.2 and v4.3 stands on its own merits and brings the Index closer to real-world problem solving, adds more private test sets to prevent gaming, and reduces saturation Intelligence Index v4.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested. Because we use a held-out test set for AutomationBench-AA, in collaboration with @zapier, the weight assigned to evaluations with private tasks or answers increases from 40% to 45%. Category weights are unchanged from v4.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20% Detailed changes: ➤ Upgraded Terminal-Bench 2.1 to 4.0: 66 multi-step tasks testing agents on tasks run in agent sandboxes driven via the terminal, including tasks involving software engineering, machine learning, science, and operations. The 4.0 update recalibrates compute and time allowances, and improves task instructions and verification. We have changed from the Terminus 2 harness to mini-SWE-agent, a minimal, model-agnostic harness. We will also be updating our Coding Agent Index, where we test model and harness pairs, to include Terminal-Bench 4.0 soon ➤ Replaced 𝜏³-Banking with AutomationBench-AA: Our implementation of Zapier’s AutomationBench tests agents on 657 business workflows across simulated applications such as Gmail, Slack, Salesforce, and Jira. Agents must complete task objectives while following business rules. AutomationBench-AA uses Zapier’s private set of 657 tasks, and is built on v1.0.6 Key results: ➤ Claude Fable 5.1 and GPT-6 Astra lead the Intelligence Index: Both Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) score 53 on Intelligence Index v4.3, followed by Claude Opus 5 (max, 51), Claude Fable 5 (with fallback, 50), Muse Spark 1.3 (max, 48) and GPT-5.6 Sol (max, 47) ➤ GLM-5.3 and Kimi K3 continue to lead open weights models (both at 44): GLM-5.3-Flash (42) is the third strongest open weights model, followed by Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (max, 36) ➤ 4 labs occupy the Intelligence vs. Cost per Task Pareto frontier: OpenAI occupies the majority of the cost-efficiency frontier, with all five reasoning efforts of the recently released GPT-6 Astra offering the lowest Cost per Task at their respective levels of intelligence. Claude Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42) and MiMo-V2.5-Pro (26) round out the rest of the frontier
1
193
Had a fun time geeking out and building with @dN0t and @devopsjacquie - using open models for coding -- they are surprisingly good (and cheaper) - using Hermes to find vintage auto parts! - new models in @nebiustf Get started by singing up to the Builder Program : lnkd.in/gkUXHpMS And share with us what you build. nitter.net/i/broadcasts/1rGmqpNjX…
1
6
152
Sujee Maniyam retweeted
The Nebius AI Builder Program is now available. AI isn’t just a model you call anymore. It’s a system you build. And builders, not a few closed labs, will decide what it becomes. The open ecosystem has all the pieces. We want to make it easier to put them together and start building. The program is free, with $400+ in credits and discounts, working code and cookbooks, office hours with engineers, and a community to build with. We’re joined by @NVIDIAAI, @LangChain, @huggingface, @cognition, @OpenHandsDev, @tavilyai, @TolokaAI, @composio, @PrimeIntellect, @MiniMax_AI, @Alibaba_Qwen, and more joining soon. Join with the link in the comments 👇
476
391
1,292
2,593,215
Here’s a quick explainer on fine-tuning, distillation, and quantization - using LEGO® bricks 🧱 🟢 Fine-tuning: specialize a pretrained model with new data. 🟠 Distillation: train a student model to learn from a teacher. 🟡 Quantization: use lower-precision weights to reduce memory and bandwidth use. Video below 👇 More detail: sujee.dev/post/fine-tuning-d… What should I explain next with LEGO: LoRA, pruning, or speculative decoding? Drop your vote in the comments 👇
1
1
6
219
Went to buy a couple of 4TB portable SSDs and almost fell off my chair looking at the prices 😳 Damn. See screenshots 👇 @demian_ai has been writing about the memory/storage crunch driven by the AI boom. It is finally sinking in for me :-) A few of his posts: - nitter.net/demian_ai/status/20644… - nitter.net/demian_ai/status/20092… - nitter.net/demian_ai/status/20766…
$PENG turned scarcity into revenue. Samsung made memory stretch. Micron locked down wafers. Korea broke the wrapper. this week, the book looked flat but the stack was being repriced ↓ Under a flat headline, the market made a clear choice: pay for what can turn scarcity into revenue now. Storage, networking, AI cloud and cooling led. HBM / packaging and power / grid lagged. $PENG, Samsung, Micron and Korea explain the rotation. REVENUE NOW - $PENG Penguin was the receipt. Q3 revenue grew 48%. Integrated Memory more than doubled. Operating income rose more than 5x. The point is not just that memory is tight. Penguin gets paid to turn memory, clusters and infrastructure software into a working AI system. Scarcity matters but making it usable pays sooner. MEMORY ELASTICITY - SAMSUNG Samsung's Blackwell test added a 1TB CXL pool for KV-cache offload. When local DRAM filled, throughput collapsed. The CXL-backed system kept running near DRAM speed (approximately 92% of DRAM performance in multi-GPU configurations). Not because CXL replaces HBM but because not every byte needs the fastest, most expensive tier. The hierarchy is widening: HBM → DRAM → pooled CXL → storage The caveat: this required host changes, a custom in-house kernel and modifications to the LMCache stack. Engineering proof, not plug-and-play adoption. But it shows where the next memory trade may form: deciding where each byte belongs. SUPPLY SECURITY - MICRON Micron put $500M behind GlobalWafers' 300mm Texas plant and signed a 10-year supply agreement. That is more revealing than another demand forecast. When the customer finances the supplier, the bottleneck has moved from the presentation deck to capital allocation. AI memory profits are being recycled into the layer beneath memory: wafers, materials and geographic resilience. RACK TIME VS GRID TIME Transformer queues now stretch beyond three years. High-voltage breakers are not far behind. Yet cooling rallied while Power & Grid fell. The physical constraint did not vanish. The market separated two clocks: Cooling gets installed with the rack. Grid equipment gets paid after permitting, financing and construction. Same density problem. Very different route to cash. DEPLOYABILITY BEAT DESTINY Networking and retimers outperformed photonics / CPO. $ANET led; $CRDO gained. That is not a verdict against optics. CPO may still be the architectural destination. But copper, retimers and systems already shipping can earn during the transition. The market paid for what can deploy before the perfect end state arrives. SAME ASSET, DIFFERENT PLUMBING SK hynix's U.S. ADR jumped on debut. The next trading day, its Seoul shares fell more than 15%, Samsung fell sharply, the KOSPI fell ~9% and trading halted. Reuters calculated a roughly 37% ADR premium after the rout. Same company, same HBM exposure, but different access, liquidity, leverage and flows. Korea did not prove that the HBM thesis was broken. It proved that the wrapper can overpower the asset. THE THREE CLOCKS 1. Revenue now: storage, networking, AI operations, cooling. 2. Scarcity later: power, fabs, packaging, substrates. 3. Elasticity: CXL, retimers, liquid cooling, orchestration. That is the map for this week: $PENG = turn scarcity into revenue. Samsung = make scarce memory go further. $MU = secure the layer underneath. Korea = price the instrument, not only the asset. The bottleneck did not disappear. The market started asking a harder question: How long until it becomes cash? Full weekly: aibottlenecks.app/alpha/week…
1
2
8
3,134
been using glm-5.3-flash on @nebiustf with @opencode for the past few days. it has been a delight to use. super fast and very smart. Handled all tasks i gave it pretty comfortably
Artificial Analysis currently ranks Nebius first among 12 providers for GLM‑5.3‑Flash on: ⚡ Output speed: ~290 tokens/second ⏱️ End-to-end response time: ~9.1 seconds A good reminder that choosing the model is only half the job. The inference stack determines the experience users actually get. See the full comparison: artificialanalysis.ai/models…
114
Sujee Maniyam retweeted
So much excitement about GLM-5.3-Flash, Kimi K3, and the pace of open models in general. But many of you are asking questions such as: does it beat what's serving in prod, will the tool calls parse in my agent loop, does latency hold when the context window fills up, can I post train it on my traces, and what's the $/token at real traffic that's what we'll cover at Builders & Brews, plus a room full of people building on open models who are worth meeting. Find your city: luma.com/g7vslxuv
Ox Alpha has a name: GLM-5.3-Flash. @Zai_org’s new open-source model is now live on Nebius Token Factory on Day 0, with a 1M-token context window, multimodal capabilities, and zero data retention. Build with it: tokenfactory.nebius.com/endp…
2
20
2,029
1️⃣ First graphic: GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index, while Kimi K3 scores 60. But look at the tradeoff 👀 GLM-5.3-Flash is: - much cheaper - farther left on the x-axis - much smaller - smaller bubble size - while staying surprisingly close in intelligence 2️⃣ Second pic: GLM-5.3-Flash running inside OpenCode. Amazing speed and performance! ⚡ Try it for yourself : tokenfactory.nebius.com/play… Model visualizer: sujee.github.io/practical-ll…
Ox Alpha has a name: GLM-5.3-Flash. @Zai_org’s new open-source model is now live on Nebius Token Factory on Day 0, with a 1M-token context window, multimodal capabilities, and zero data retention. Build with it: tokenfactory.nebius.com/endp…
1
1
233
I tested 3 NVIDIA Nemotron models on @nebiustf with the same sales-analysis agent. All scored 100%. Nemotron 3.5 Lightning finished in 7s for $0.0019 - about 7x cheaper than Super and 16x cheaper than Ultra. TLDR; For focused agent workflows, smaller may be enough. piped.video/watch?v=Ycmew5kJ… Code: github.com/nebius/token-fact…
1
5
131
Looking fwd to coding with GLM 5.3 flash on TF!
OK Now officially Very soon at @nebiustf The team is on that:)
2
162