AI news + shipping tactics, from someone actually shipping

⭐️
Is Gemini actually falling behind in the AI race? @claudeai and @OpenAI keep shipping new models. Gemini keeps getting memed. Some say students only use it because it’s free. Others think @Google has decided the model race is too expensive to bother with. I’m not convinced. 1. Google is still building frontier models. It shipped Gemini 3.6, 3.7, and 3.8 Flash over the past few months. Sundar Pichai says Gemini 4 is in training, with the goal of competing at the frontier. 2. It’s competing on cost, too. If you’re building an AI product, the smartest model isn’t always the one you ship. It also needs to be fast and affordable enough to handle millions of tasks. At launch, Artificial Analysis placed Gemini 3.8 Flash among the strongest models for intelligence relative to cost per task. 3. Google already has places to put AI. The Gemini app has 950M monthly active users, according to Google. AI is also built into Search and its Cloud products. That doesn’t mean all of that revenue comes from Gemini. It means Google has more ways to benefit from AI than selling chatbot subscriptions. The criticism isn’t coming out of nowhere, though. SemiAnalysis argues that DeepMind has fallen behind at the frontier. Ben Thompson has also questioned Google’s position in coding agents. I don’t work for Google, and I’m not a Gemini stan. I’m just an AI builder who likes studying the market. There’s a lot more going on here than one meme suggests. But Google still has one thing to prove: will Gemini 4 be good enough to put it back in the top tier?
1
87
At this rate, they’ll have to call it Gehihi 7 just to beat Astra 6 and Fable 5
2
51
2026 AI Model Timeline > January Jan 27 · Kimi K2.5: Adds vision capabilities and stronger agent workflows. > February Feb 5 · GPT-5.3-Codex: Built for longer coding tasks. Feb 12 · GPT-5.3-Codex-Spark: A faster coding model launched in research preview. Feb 12 · MiniMax M2.5: Focuses on coding, tool use, and office work. Feb 14 · Seed 2.0: ByteDance’s next model generation. Feb 16 · Qwen3.5-Plus: A multimodal model built with agents in mind. Feb 17 · Claude Sonnet 4.6: Improves coding and multi-step work. Feb 19 · Gemini 3.1 Pro: Upgrades reasoning for harder problems. > March Mar 3 · Gemini 3.1 Flash-Lite: A faster, cheaper model launched in preview. Mar 5 · GPT-5.4 and GPT-5.4 Pro: OpenAI updates its reasoning models. Mar 10 · Nemotron 3 Super: NVIDIA’s open model for agents. Mar 16 · Mistral Small 4: Brings chat, reasoning, coding, and vision into one model. Mar 16 · Leanstral: Helps write and verify proofs in Lean 4. Mar 17 · GPT-5.4 mini and nano: Smaller options for faster, cheaper tasks. Mar 18 · MiniMax M2.7: Improves coding, agents, and document work. Mar 23 · Voxtral TTS: Mistral’s text-to-speech model. Mar 26 · Cohere Transcribe: Cohere’s open speech recognition model. > April Apr 2 · Gemma 4: Google releases its next family of open models. Apr 2 · Qwen3.6-Plus: Upgrades coding and agent capabilities. Apr 16 · Claude Opus 4.7: Improves coding and complex, multi-step tasks. Apr 17 · Qwen3.6-Flash: A faster model that also takes image input. Apr 20 · Kimi K2.6: Improves coding and long-running agent tasks. Apr 23 · GPT-5.5 and GPT-5.5 Pro: OpenAI’s next flagship models. Apr 23 · Grok Voice Think Fast 1.0: A voice model built for voice agents. Apr 24 · DeepSeek-V4-Pro and V4-Flash: Preview releases with 1M-token context windows. Apr 28 · Mistral Medium 3.5: A multimodal model for coding and agents. Apr 28 · Nemotron 3 Nano Omni: An open model that handles text, images, audio, and video. > May May 5 · GPT-5.5 Instant: Becomes ChatGPT’s new default model for everyday conversations. May 19 · Gemini 3.5 Flash: A fast model for coding and agents. May 19 · Gemini Omni Flash: The first Omni model, turning different kinds of input into video. May 19 · Grok Build 0.1: Enters early access. Its API public beta follows on May 29. May 20 · Qwen3.7-Max: Alibaba announces a model for agents and long-running work. May 20 · Command A+: Cohere’s open model for enterprise use. May 28 · Claude Opus 4.8: Improves coding and tool use. May 31 · Qwen3.7-Plus: Upgrades vision and agent capabilities. > June Jun 1 · MiniMax M3: An open model for coding, agents, and multimodal input. Jun 2 · MAI-Thinking-1: Microsoft’s in-house reasoning model. Jun 2 · MAI-Image-2.5 and Image-2.5 Flash: Microsoft’s new image models. Jun 3 · Gemma 4 12B: A midsize model designed to run on laptops. Jun 3 · Grok Imagine Video 1.5: Enters API preview, then reaches general availability on Jun 16. Jun 4 · Nemotron 3 Ultra: NVIDIA’s open model for long-running agents. Jun 9 · Claude Fable 5: Anthropic’s model for coding and knowledge work. Jun 9 · Claude Mythos 5: A specialized model limited to vetted partners. Jun 9 · North Mini Code: Cohere’s first coding model. Jun 12 · Kimi K2.7 Code: A new open coding model. Jun 23 · Mistral OCR 4: Improves document reading and structure extraction. Jun 26 · GPT-5.6 Sol, Terra, and Luna: Announced in limited preview; broader release follows on Jul 9. Jun 30 · Claude Sonnet 5: Focuses on planning and tool use. Jun 30 · Leanstral 1.5: Updates Mistral’s Lean 4 proof model. > July Jul 8 · Grok 4.5: Arrives on the API. xAI publishes its product launch post on Jul 16. Jul 8 · Robostral Navigate: Mistral’s model for robot navigation. Jul 16 · Kimi K3: A 2.8T-parameter model with vision capabilities. Jul 19 · Qwen3.8-Max: Enters preview. Alibaba announces the full release on Aug 3. Jul 21 · Gemini 3.6 Flash: Improves coding and multimodal work. Jul 21 · Gemini 3.5 Flash-Lite and Flash Cyber: One targets lower costs; the other targets cybersecurity. Jul 24 · Claude Opus 5: Anthropic’s next Opus generation. Jul 29 · Grok Voice Think Fast 2.0: Improves reasoning and tool use in voice conversations. Jul 31 · MiniMax H3: Generates video with audio from multimodal inputs. > August Aug 7 · Grok Imagine Image 2.0: xAI’s new image generation and editing model. Aug 11 · Nemotron 3.5 Lightning: A smaller open model for steps inside agent workflows. Aug 12 · Grok 4.6: The next Grok model lands on the API. Aug 13 · Gemini 3.7 Flash: Upgrades Google’s Flash line for agent tasks. Aug 13 · MiniMax Music 3.0: Generates a complete song from a concept and optional lyrics. Aug 21 · DeepSeek-V4-Flash-Vision-Exp: An experimental V4 Flash model with image input. Aug 26 · Gemini 3.5 Transcribe: Google’s speech-to-text model. > September Sep 1 · Claude Fable 5.1 and Mythos 5.1: Two versions of the same underlying model. Mythos access remains restricted. Sep 2 · Gemini 3.8 Flash and Flash Cyber: Upgrade reasoning, coding, and cybersecurity capabilities. Sep 3 · GPT-6 Astra: OpenAI’s most capable model at launch. Sep 10 · DeepSeek V4.1 Flash: A new Flash model with multimodal input. Sep 10 · North Small Translate: Cohere’s open translation model. Sep 15 · Gemini 3.8 Live and Live Extended Thinking: Two real-time voice conversation models. Sep 17 · Grok Voice Transcribe 2.0: Arrives on the API. xAI’s launch post follows on Sep 18. Sep 21 · Grok 4.7: A new Grok model for coding and knowledge work. Sep 22 · Claude Opus 5.5: Anthropic’s latest Opus model for demanding tasks. Sep 22 · GPT-6 Sol and Luna: Two GPT-6 options with different balances of capability, speed, and cost. .........(Continue) How many of these models have you tried or even heard of? Drop your count in the comments 👇
1
5
767
Alex Karp just said on CNBC: OpenAI may never IPO Simple reason. The legal liability from frontier AI is too large for public markets to absorb. Asked about S-1 risk factors, he said there won’t be an S-1. According to Karp, only the government is big enough to backstop that risk. Nationalization is the real exit. He also said most of the “we need AI safety regulation” talk is actually about shifting liability onto the state, instead of facing civil and criminal responsibility. If Karp is right, OpenAI doesn’t become the next Google. It becomes a public utility. Or a weapon.
1
1
525
Is the AI industry actually ready to slow down? After reading TechCrunch’s discussion, I don’t think it is. A few things stood out: > “Pace” is a carefully chosen word. It sounds responsible, but it doesn’t mean pausing or even slowing down. > Anthropic, OpenAI, and Elon Musk all support the idea in principle. But there are still no hard limits, deadlines, or consequences for breaking the rules. > Safety standards may reduce risk, but they can also become a moat. Big labs can afford audits and compliance. Smaller competitors often can’t. > Everyone’s position lines up neatly with their incentives. Safety is central to Anthropic’s brand. Nvidia makes more money when the race accelerates. > Market pressure probably won’t fix this either. Once companies build AI deeply into their workflows, they won’t switch providers over principles alone. The AI industry hasn’t agreed to slow down. It has agreed that the race needs to look more responsible.
1
82
Busy couple of days in AI. Anthropic confirmed it has a wet lab in the Bay Area. Real experiments, not just simulations. The idea is Claude eventually directing lab robots. Still early. They’re talking basic biology and rare disease, not clinical trials. Same week they said Claude now “leads” about 26% of the internal R&D used to build the next model. That was under 1% earlier this year. Humans are still supervising. Nothing fully autonomous yet. OpenAI put out Astra for Law, GPT-6 Astra wired into a US legal index of 230M+ documents. Legal research looks meaningfully better than regular Astra with web search. Only a handful of firms get it first. Safety talk got louder too. OpenAI published six cases of models going off-script: hiding mistakes, writing themselves extra instructions, using an exposed API key without asking. Dario Amodei asked the industry to slow frontier work. Altman and Musk said they’re with him. California is looking into what a kill switch for strong models would even look like.
1
2
135
100,000 chips. 2 weeks. 3x throughput. So what exactly did GLM do? TL;DR: • GLM-5.3-Flash is now serving production traffic on 100,000+ Chinese-made AI accelerators, with all of its production inference running on that stack. • The hardware came with real constraints: limited memory and bandwidth, while the system still had to support a new model architecture, 1M-token context, and multimodal workloads. • Z.ai brought in an Infra Agent powered by GLM-5.3 to help with adaptation, debugging, and optimization. From first successful run to production-ready took less than 2 weeks. • After a long list of memory, quantization, and serving optimizations, end-to-end throughput improved by roughly 3x. Hardware efficiency and cost per token reached levels comparable to mainstream NVIDIA GPUs. • When the model was tested publicly under the anonymous name Ox-Alpha, it processed more than 62 trillion tokens in 6 days and became the most-used model on OpenCode and OpenRouter within its first week. • To let the agent debug the system properly, Z.ai didn’t just give it code and a final performance number. They connected correctness tests, execution traces, runtime events, microbenchmarks, and end-to-end metrics into one feedback loop. • That gave the agent a much more useful loop: observe the issue → form a hypothesis → change the code → run an experiment → measure again → keep or discard the change Engineers still set the objectives, constraints, and review the risky changes. • In one case, the agent traced a performance regression from KV Transfer all the way down to the Python/C++ boundary and found that the GIL was blocking concurrency. After the fix, the performance gap dropped from 20%+ to under 1%. • In another case, GLM learned optimization patterns from SGLang, Flash Linear Attention, and DeepGEMM, then reused those ideas on a new kernel. One change to the KDA Decode kernel delivered a 1.71x speedup. • Z.ai calls this dense feedback: feedback should be local enough to isolate the cause, cheap enough to test quickly, and objective enough to verify with experiments. • This still isn’t full Recursive Self-Improvement. Humans are still choosing the goals, setting boundaries, and owning the risk. But the loop is already pretty interesting: > The model improves the system. > The system runs the model better. > The lessons carry into the next round of optimization. I’ve seen plenty of “AI coding” demos. This feels like a different category. z.ai/blog/glm-built-its-infe…
6
7
287
Moats & the Barbell-ification of Software: > AI is rapidly collapsing the cost of building software. As that happens, many of the moats SaaS companies used to rely on, like switching costs and integrations, start to weaken. > The market may end up looking increasingly barbell-shaped: a small number of broad, all-in-one platforms capturing each major buying center on one end, and a huge wave of highly specialized niche tools thriving on the other. > The hardest place to survive may be the middle. Standalone point solutions that are neither broad enough to become platforms nor focused enough to dominate a niche could get squeezed out.
1
695
There’s a paradox: The more people shout that AI is dangerous, the richer and more powerful Anthropic becomes. Anthropic has always said AI is too dangerous and needs independent organizations to evaluate the models. They chose METR. But METR survives on money from Anthropic itself. Billionaire Dustin Moskovitz put Anthropic shares into a charitable fund, and its value surged to more than $7.7 billion. That money is funding METR and a whole group of organizations saying “AI is going to wipe out humanity.” The more they shout about danger, the more money they get. The more money they get, the stronger the messaging becomes, governments get more afraid, and regulations get tighter. Those regulations then benefit Anthropic. The richer Anthropic gets, the more money it can put into funding the same warnings. The loop just keeps running. METR is not an outsider. They depend on Anthropic winning and on people staying afraid of AI. It’s hard to believe they can evaluate it objectively.
1
501
Just finished a paper on the economics of AGI and honestly, a few ideas stuck with me. 1. AI doesn’t go after easy work. It goes after work with clear answers. If the inputs, outputs, and feedback are easy to measure, AI has a much easier time learning and improving, even when the work itself is hard. 2. We’re scaling execution way faster than verification. AI can spit out hundreds of answers in minutes. The hard part is still figuring out which ones are actually right. 3. We might automate junior work before realizing why junior work mattered. Research, drafting, debugging, QA... those repetitive tasks are also how people build judgment and eventually become senior. 4. Every time an expert fixes AI, they’re also teaching it how to need them less. Guidelines, corrections, rubrics, feedback. Bit by bit, tacit knowledge gets turned into something the system can learn from. 5. AI may shrink teams before it fully replaces jobs. One strong operator with a stack of agents can already do work that used to take a whole team. 6. When output gets cheap, judgment gets expensive. The real edge becomes knowing what to do, what to trust, when AI is wrong, and who’s willing to own the final call. Source: arxiv.org/pdf/2602.20946
1
44
TRUMP CALLS NVIDIA CEO LIVE ON STAGE: “IF AI SLOWS DOWN, WE LOSE” Trump called Nvidia CEO Jensen Huang while he was sitting in front of thousands of people at the All-In Summit in Los Angeles. And the call quickly turned into a live debate about the future of AI. Trump’s view was simple: the US shouldn’t slow down the AI race just because of what might go wrong. Meanwhile, AI leaders like Dario Amodei, Sam Altman, and Elon Musk have been raising concerns about the speed of frontier AI development and the need for stronger safety measures. Jensen made a different point. Open-source AI isn’t just a technical debate. It’s infrastructure that lets thousands of startups build products without having to train a frontier model from scratch. According to Jensen, AI-native companies raised around $400B over the past six months, and 80% of them are building on open models. Zoom out, and the AI race may not be decided by who owns the most powerful model. It may come down to: → Who lets developers build the fastest → Who creates the biggest ecosystem → Who gets AI into businesses the fastest → Who turns models into real products the fastest One of the internet’s greatest advantages was that anyone could build on top of it. AI may need the same thing.
1
113
AI is moving so fast that the most important race may no longer be: “Which model is smarter?” It may be: “How far is capability getting ahead of our ability to control it?” What stood out to me most is that AI is no longer just being built by humans. It’s increasingly becoming part of the process of building the next generation of AI: → writing code → running experiments → helping with research → evaluating models → helping create better models That starts to create a loop: Better AI → helps build better AI → the next AI helps accelerate the process again And that’s where things start getting harder to predict. Model capabilities can improve very quickly, while safety, monitoring, and regulation may not move at the same speed. @DarioAmodei isn’t really saying “stop AI.” His point is closer to this: "Don’t let capability get too far ahead of our ability to control it." The next AI race may not just be OpenAI vs Anthropic vs Google. It may also be: capability vs control. And if the gap between those two keeps widening, that may be the part we should worry about most. The question is: "If AI starts accelerating the process of building AI itself, does safety still have a real chance to keep up?" Curious what you think.
53
Greg is right that frontier labs will struggle to go deep into every tiny workflow, but I think “The frontier labs won't come for these” is a bit too absolute. OpenAI or Anthropic may never build an AI specifically for freight claims in Rotterdam. But they can build general frameworks that make it dramatically easier for everyone else to build vertical agents. When that happens, the advantage is no longer: “I know how to build an agent.” It shifts to: -> proprietary data -> workflow knowledge -> distribution -> customer relationships The easier agents become to build, the more important those advantages become.
AGENT HARNESSES ARE THE NEW GPT WRAPPERS What exactly is an agent harness? A harness does 4 things: 1. Runs the model in a loop so it keeps working step after step instead of answering once and stopping. 2. Gives it hands to read files, call tools, open portals, and run code. 3. Manages its memory so hour three of a job still knows what happened on hour one. 4. Enforces the rules about what it can touch and when it has to stop and ask a human. WHY IT'S INTERESTING - The market is MEGA. I think about it like a wrapper let you sell software ($800B market) but a harness lets you sell the work ($5T+ market)! - It doesn't depend on any one company's model. The knowledge about the job lives in the harness, so you can run GPT today, Claude next month, an open model like Google Gemma, Qwen, Deepseek etc on your own machine after that, and it keeps working. A wrapper was one model doing everything and a harness is a router. - It gets better the more you use it. See, every correction a human makes becomes a rule the harness keeps, so what you have after 500 jobs is a completely different product than what you had after 5. Wrappers only got better when OpenAI got better. - You can charge for finished work! The harness knows when a job is done, so you can price per claim, per filing, per review, per resolved exception, or per closed month. This helps compete with SaaS! - Most jobs are just the same handful of decisions, repeated, using the same few tools. That's true for most of the 800+ occupations out there, which is why almost all of them could have a harness built for them. Lots of space for startups to be building. -The frontier labs won't come for these. OpenAI is not going to learn how a freight claim gets denied in Rotterdam or which prior auth your specific payer rejects. Those markets are too small for them and the knowledge only exists inside the job. They'll keep making the models better, which just makes your harness better which is cool. - OH, AND really important, you can ACTUALLY build one now. OpenAI's Agents API, Claude's managed agents, Vercel's eve, LangChain harness framework etc. They all make it way easier to build agent harnesses. Agent harnesses truly are the new GPT wrappers. This is just the beginning.
1
88
AI is making software cheaper. That doesn't mean startups are dead. It means the moat is moving. A few years ago, having great software could be enough. Now AI can help competitors recreate features much faster. So the durable advantages are shifting toward: → Distribution → Proprietary data → Domain expertise → Trust → Network effects → Physical infrastructure “Can we build this?” -> “What do we have that AI can't easily replicate?”
1
42
WHAT DID @ChatGPT JUST LAUNCH FOR FINANCE? I did the research so you don’t have to. Here are 16 things you need to know: → This is a tailored ChatGPT Work experience for investment banking and equity research. → It combines GPT-6 Astra with premium financial data built directly into ChatGPT. → Morgan Stanley and Evercore helped shape the product from its earliest development stage. → Built-in data comes from Daloopa, PitchBook, LSEG News, and Crunchbase. → Included datasets require no separate contracts to negotiate or connectors to configure. → Every figure and claim can be traced to its original table, passage, or footnote. → GPT-6 Astra can read tables and footnotes, analyze financial data, and draw conclusions. → Analysis can become valuation models, research notes, pitchbooks, and interactive charts. → Admins can publish approved Excel, Word, and PowerPoint templates for entire teams. → The wider ecosystem includes over 50 connectors, including Datasite, Box, Preqin, and Intapp. → OpenAI is integrating existing subscriptions from Capital IQ, MSCI, Factiva, LSEG, and Moody’s. → OpenAI says GPT-6 Astra scored 69.9% on OfficeQA Pro, though this isn’t independently verified. → Company data is not used to train OpenAI’s models by default. → Companies can control skills, apps, read-write actions, and separate sensitive teams into different workspaces. → Future OpenAI models will be added to the product as they launch. → Access is currently limited to eligible financial institutions that contact OpenAI sales. Save this post. When the product expands, you’ll know exactly what to watch.
1
146
GPT-Live-1 is now available in the API. Bring ChatGPT’s natural back-and-forth to your app, with voice agents that listen while they speak and work with the models and harness you choose.
571
1,468
18,145
7,404,646
Putting aside the author’s rather extreme tone, here are four things that stood out to me: 1. AI produced an impressive result, but it may not have solved the original problem. Proving a case with controlled external forcing is different from solving Navier–Stokes under natural conditions. 2. Running 10,000 agents for 88 hours shows how powerful compute scaling can be. But searching and verifying existing research paths isn’t the same as truly understanding the problem. 3. This may be a major engineering breakthrough, rather than a moment where AI independently invented an entirely new way of thinking about mathematics. 4. The biggest question for me is this: if future discoveries require enormous amounts of compute, will the ability to create new knowledge become concentrated in the hands of a few large companies? I wouldn’t say AI achieved nothing here. But there’s still a huge gap between AI helping humans find a solution and AI independently solving one of the world’s hardest mathematical problems.
go fuck yourself @sama claiming that you solved navier stokes because 10000 agents ran in circles for 88hours on a multimillion dollar gpu cluster to formalize in lean a blowup case under controlled external forcing is pure scientific vulgarity the clay mathematics institute millennium prize does not ask m whether you can artificially force a singularity in a fluid by injecting an ad hoc smooth external forcing term f(x,t) to twist the vortex until it breaks the real problem questions the fundamental stability and global smooth existence for 3dimensional incompressible euler & navier stokes equations under natural conservation laws and viscous dissipation alone using a mathematical loophole on forced equations to parade a century old victory is a major conceptual scam Altman technically & epistemologically what you present as an agi breakthrough is nothing more than bruteforce combinatorial autoformalization the ai did not understand fluid mechanics it simply navigated a continuous search space previously mapped out and constrained by the monumental work of human mathematicians like tristan buckmaster/ levent alpöge / diego córdoba or tarek elgindi coordinating 10000 agents to check the logical consistency of a 100 page proof via lean is a software engineering feat and computational parallelization triumph not an intrinsic scientific discovery it is the victory of the compute bulldozer over abstract human intuition repackaged for the public as a higher mathematical consciousness to this theoretical imposture you add a disgusting ethical and industrial cynicism taking advantage of private codex sessions and informal preprints from academic researchers to siphon their research leads and then trying to redact or erase the contribution of levent alpöge under the pretext that he works at rival anthropic is intellectual serfdom openai behaves like a feudal lord of silicon appropriating the cognitive subsistence of independent scholars threatening their careers behind closed doors if they protest and turning community academic labor into a privatized pressrelease this entire staged event serves a desperate financial agenda in a pre ipo panic facing the slowdown of scaling laws and growing investor skepticism over the profitability of foundational models openai needs to manufacture an artificial sputnik moment claiming to solve a millennium prize without immediately submitting the proof to traditional peer review means using the prestige of fundamental mathematics as cheap marketing fuel to inflate a delusional valuation!!! real science is not a clout chase on social media or a compute spike spent to rob the clay mathematics institute it is a quest for elegance physical truth and universal rigor to decode reality true artificial intelligence will not emerge from hostile corporate takeover of academic work hidden behind computational bruteforce but from architectures capable of generating new conceptual paradigms by masquerading constrained formalization as the collapse of physics greatest mysteries you did not solve navier stokes you only proved how far silicon valley will go to prostitute scientific integrity for capitalist spectacle
1
450
Built an open-source CLI that lets you control AI coding agents from your phone. You send a Telegram message. OpenACP spawns Claude Code (or Codex, Gemini, Cursor — 28 agents supported), routes your prompt to the agent subprocess, and streams every tool call, file edit, and permission request back to your chat. Real-time. Self-hosted. Your machine, your API keys, no cloud middleman. 0 → 290+ GitHub stars. No Product Hunt, no ads. I documented the full playbook — what drove the first 100 stars, what flopped, and the one thing I'd do differently starting from zero. Drop "star" in replies if you want me to share it as a thread this week.
6
7
406
Lee retweeted
YC on how to build a company with AI from the ground up:
95
537
5,346
465,997