Growth Building Real-World AI Infra Ex-NEAR | Ex-VC

Every builder who uploaded source code to Mythos just learned the same lesson: the model provider owns the relationship. ZDR changes without notice, the stack shift happens quietly. Is the new vendor lock-in proprietary dependency? Build accordingly.
Fable 5 is a Trojan horse. I thought everyone saw it. After a bunch of conversations with CTOs and CEOs this week, I realized many didn't. Anthropic has been one of the best platform partners in AI. They understood something early that others missed: nobody is going to build serious enterprise software on top of you without trust. Audited ZDR, strong enterprise controls, clear boundaries around customer data. That trust paid off and we built on them because of it. At the same time, this is a brutally competitive market. Anthropic has great models, great marketing, strong distribution, and enormous mindshare. But there is no world where model performance stays differentiated forever. Token costs will fall. Inference gets optimized. New models arrive. The model layer is a knife fight. So where do you go next? You move up the stack. They didn't hide this at all and they started working on it. They realized it’s a hard problem. Business workflows, the software up the stack isn’t trivial. Our APIs aren’t text in, text out. We have build businesses around the messy way in which the world works. I think they realized that it's harder than they thought. Now enter Mythos. They launched it as a too good to share because it finds vulnerabilities in your code. The pitch is almost impossible for a CTO to ignore: point it at your codebase and it will find vulnerabilities, security issues, and bugs that your team missed. They beat that drum for a few months. They even leaked some of the findings to the large tech companies, which would pull the smaller ones into security war-rooms with them to make plans on how to use Mythos when it launches. The frenzy was building. The first thing most software companies will do when it launches is exactly that. They won't upload customer data. They'll upload their source code. Their crown jewels. Then Fable 5 is released, ZDR is altered, what do you think every CTO did? Insane.
2
3
344
40+ models in six months. The model layer is commoditizing in real time. None of this is a moat. The value moves to whoever deploys it into a real workflow. Most people stay loyal to three anyway.
HERE’S A LIST OF EVERY AI MODEL LAUNCHED THIS YEAR SO FAR: • Qwen3-Max-Thinking • Kimi K2.5 • Step-3.5-Flash • Claude Opus 4.6 • GPT-5.3 Codex • GLM-5 • Claude Sonnet 4.6 • Param-2 • Sarvam-105B • Sarvam-30B • Gemini 3.1 Pro • GPT-5.4 • Mistral Small 4 • MiMo-V2-Pro • Gemma 4 • GLM-5.1 • Muse Spark • Qwen3.6-35B-A3B • Claude Opus 4.7 • GPT-5.5 • DeepSeek-V4-Flash • DeepSeek-V4-Pro • MiMo-V2.5-Pro • MiMo-V2.5 • Gemini 3.5 Flash • Claude Opus 4.8 • Step 3.7 Flash • GPT-5.5 Instant • Grok 4.3 • Granite 4.1 30B • Qwen3.7 Max • MiniCPM5-1B • JT-35B-Flash • MiniCPM-V 4.6 1.3B • Ring-2.6-1T • MiniMax-M3 • Qwen3.7 Plus • Gemma 4 12B • Nemotron 3 Ultra 550B A55B • DiffusionGemma 26B-A4B • Kimi K2.7 Code • Claude Fable 5 • MAI-Code-1-Flash • MAI-Thinking-1 • U2 We’re only half way through the year..
3
261
Validating what Im seeing. Many teams are racing to fix cost and deployment, not the model. 98.4% harness, 1.6% model. As models converge, the moat moves to the scaffolding around them. Permission gates, recovery logic, context handling. The model reasons. The system ships.
Claude Code fully dissected! Researchers from UCL reverse-engineered the leaked Claude source. What they found changes how you should think about agent design. Only 1.6% of the codebase is AI decision logic. The other 98.4% is operational infrastructure. Permission gates, tool routing, context compaction, recovery logic, session persistence. The model reasons. The harness does everything else. This is the opposite of what most agent frameworks do today. LangGraph routes model outputs through explicit state machines. Devin bolts heavy planners onto operational scaffolding. Claude Code gives the model maximum decision latitude inside a rich deterministic harness, and invests all its engineering effort in that harness. The core loop is a simple while-true. Call model, run tools, repeat. But the systems around that loop are where the real design lives: A permission system with 7 modes and an ML classifier. Users approve 93% of prompts anyway, so the architecture compensates with automated layers instead of adding more warnings. A 5-layer context compaction pipeline. Each layer runs only when cheaper ones fail. Budget reduction, snip, microcompact, context collapse, auto-compact. Four extension mechanisms ordered by context cost. Hooks (zero), skills (low), plugins (medium), MCP (high). Each answers a different integration problem. Subagents return only summary text to the parent. Their full transcripts live in sidechain files. Agent teams still cost roughly 7x the tokens of a standard session. Resume does not restore session-scoped permissions. Trust is re-established every session. That friction is the point. The bet behind all of this is simple. As frontier models converge on raw coding ability, the quality of the harness becomes the differentiator, not the model. Paper: Dive into Claude Code (arXiv:2604.14228) We've shared an article on Agent Harness and what every big company is building. Read it below.
3
205
This is the whole game. Cars had 10M on the road feeding the loop. Humanoids had nothing. You cant train a robot that was never deployed. Sim lies about friction and slip. Real physics doesnt. The academy isnt a robot story. Its a data flywheel.
Elon Musk reveals SpaceX is building a 30,000-robot academy where humanoids learn from each other. Cars were easy. Tesla had ten million on the road, beaming back driving data every second. But humanoid robots? There weren't ten million Optimi yet. There weren't ten. Robotics had run data-starved for decades. Tesla decided to fix it. You couldn't train a humanoid that had never been deployed. So Musk built a school for them instead. "We can have at least 10,000 Optimus robots, maybe 20-30,000, that are doing self-play and testing different tasks." Tesla called it the Optimus Academy. Picture a warehouse the size of a chip fab. Thirty thousand humanoid robots inside. Picking things up. Folding clothes. Walking. Tripping. Catching themselves. Failing in ways no human roboticist had thought to script. Each watching the others, learning what the human body shouldn't have made look easy. Every move generated a data point. Every failure generated a sample. Every robot taught every other robot. In simulation, Tesla could spin up a million robots overnight. But simulated physics lied about friction, slip, and drift. Real physics didn't. Cars learned from drivers. Optimi learned from each other. Each generation made the next one cheaper, faster, smarter. By the tenth generation, no human would recognize the curriculum. Recursive learning at electromechanical scale. Musk, on closing the loop: "You use the tens of thousands of robots in the real world to close the simulation to reality gap." Whoever opened the academy first owned the species. P.S. I made a playbook breaking down 100+ most powerful decision making mental models used by history's greatest thinkers. 5,000+ downloads. 113 five-star reviews. Grab a free copy here: besuperhuman.gumroad.com/l/m… If you're new here, follow @GeniusGTX for content on the greatest minds in economics, psychology, and history. — Elon Musk ( @elonmusk ), CEO of Tesla and SpaceX, on Dwarkesh Patel's ( @dwarkesh_sp ) podcast
1
117
Compute is right, but only partly. Designing jet engines and medical devices isnt blocked by GPU access. Its blocked by materials data, test results, and failure modes that exist nowhere online. The scarce input needed for an artificial general engineer is industrial data, not just racks.
Jeff Bezos on CNBC explains revealed what Prometheus is building. Today his new company Prometheus announced a $12B funding round at a valuation of $41B . Prometheus trying to build an artificial general engineer that can help design and manufacture physical products like engines, medical devices, and electronics. So the target areas are hard physical products like jet engines, chips, bridges, medical devices, consumer electronics, aerospace systems, vehicles, and drug design, where design cycles can take years because every idea has to survive physics, materials, cost, testing, and factory limits. Bezos’ jet-engine example explains it well: asking for the same engine with 10% more thrust can become a 10-year engineering program, and Prometheus wants to shrink that “dream-build” cycle by 10x or more. The $6.2B launch funding gave Prometheus a massive starting base, and the new raise says the company likely needs far more compute, talent, and industrial data before it can prove the product. Their $41B valuation shows that frontier AI is becoming less a software race than a compute procurement race. A company with no broadly shipped product can raise $12 billion at a $41 billion valuation because investors are not only funding a model, they are prepaying for the machines that might make the model possible. The scarce asset is no longer just talent or algorithms, but clustered GPUs, power contracts, cooling, networking, and the operational skill to keep expensive silicon busy. They are proof that demand is arriving faster than infrastructure can be built, and that every frontier funding round quietly turns into a future claim on power, racks, GPUs, and uptime.
120
The headline is world models. The real bet is the last line: robot collects data, model gets better, robot gets better. Pretraining gets you to deployment. The loop after deployment is what compounds.
We’re going all in on World Models. Today we’re launching the 1X World Model Lab. The bet is simple: You can’t fine-tune your way to AGI. And you definitely can’t fine-tune your way to robots that can operate in the physical world. General-purpose humanoids need models that understand space, motion, objects, causality, affordances, physics, and action before they ever see a specific task. The frontier is not better VLA wrappers. The frontier is embodied world models. The 1X World Model Lab will focus on large-scale embodied world model pretraining: building the most generalizable foundation model for humanoid robots from the ground up. The next frontier in AI requires scaling: web-scale media + egocentric human videos + sim + dexterous remote operated robot data + on-policy NEO data → real-world deployment for robot data collection and RL → abundance of data → physical AI The robot collects data. The model gets better. The robot gets better. Repeat. To lead this, we brought in one of the best for the mission: @_sam_sinha_ , as Head of World Models. Sam was a founding research scientist at Luma AI and has been at the frontier of scaling multimodal generative video models his whole career. If you’re the best in the world at large-scale pretraining, video models, robotics, RL, infra, or data — and you want your models to move atoms, not just pixels — join us. Send background + evidence of exceptional ability to: wmlab@1x.tech We’re building the model that makes autonomous labor real.
6
239
Clear use case here. Real city setting, constrained task. Fixed dock, known location, predictable approach. Its a deployable pattern, narrow and controlled even in the real world. What happens when someone grabs the bike before it docks?
A robot just parked and docked a Citi Bike by itself in the middle of NYC. 🤖🚲 We’re entering the era where AI won’t just think… it will interact with the real world beside us. 👀 #AI #Robotics #IoT #5G #Tech #innovation #SmartCities
2
112
15-30 min sounds small. It works because the failures are in there, not just the successes. Simulation gets the task right. Real data teaches where the sim was wrong.
15–30 minutes of real-world robot data. That's now enough to go from sim-to-real failure to working robot. Let’s see… You train a robot in simulation. You deploy it in the real world. It fails. The physics don't match. So you try to fine-tune it with real data, but… you never have enough real data, and the fine-tuning breaks everything the simulation taught it. SimDist fixes this with one key decision: don't transfer the policy. Transfer the world model. Keep the reward and value knowledge from simulation frozen. Only update the part that's actually wrong, how the robot predicts physics. Now the robot doesn't have to relearn the entire task in the real world. It already knows what success looks like. It just needs to correct its understanding of how the real world moves. The part that makes this work: they also trained on failures and recoveries; not just perfect demonstrations. Without that, the planner finds the gaps and exploits them. With it, the robot can tell a good future from a bad one. That's all it needs. Results on peg insertion, table leg assembly, locomotion on slippery and uneven surfaces. Tasks that require precision, force, and quick reaction. Thanks for sharing, Tyler Westenbroek ([@ty_westenbroek]. Interactive visualization + paper: sim-dist.github.io ——- Weekly robotics and AI insights. Subscribe free: 22astronauts.com
1
3
144
LLMs were not designed for real-world action. Real world is messy and dangerous. World models are supposed to bridge that. Still a lot to figure out on how...
Yann LeCun says you cannot build a reliable agentic system without a world model LLMs don't have world models. They can't predict the consequences of their actions before taking them "they just act, and whatever happens next is someone else's problem" Without that, it's not intelligence
4
183
High-risk industrial is a strong first market for physical AI. Clear value of safety. Controlled enough to deploy reliably. Dangerous enough that the ROI is obvious. Factory → hazardous industrial → less controlled environments.
A humanoid robot is now climbing vertical industrial tanks to weld, grind, inspect, and remove rust. 🧲🤖 This is RobotPlusPlus’ embodied-intelligence wall-climbing robot, now deployed in high-risk scenes across chemical plants, shipyards, and energy facilities. The upper body uses dual humanoid arms. The lower body uses a wheeled magnetic-adhesion chassis that lets the robot work on steel walls instead of flat factory floors. The reported specs are industrial, not demo-stage: 90 kg body weight, 15 degrees of freedom, 12 active arm joints, millisecond-level remote response, and tethered power for long-duration operation. The real advantage is tool switching. Swap the end effectors, and the same platform can move between welding, grinding, flaw detection, rust removal, spraying, and surface treatment. For operators, the workflow changes from climbing scaffolds or hanging in baskets to controlling the robot through a remote interface and VR glasses. RobotPlusPlus says its special-operation model has learned from 100,000+ hours of field work, 22,500 km of operating distance, and 5,000 km² of covered work area. This is where embodied AI starts to look useful: not dancing onstage, but taking tools into places humans should not have to enter.
2
122
50+ builders at @fdotinc SF lab for the Physical AI Hack World Tour with @Ryan_Resolution @IsaacSin12 @seeedstudio The highlight: watching 12 teams train robotic arms with 40-50 video demonstrations over 2-5 hours to complete physical tasks.
4
9
53
5,890
SF stop of the Physical AI Hackathon Tour Excited for this!
Glad to partner with @markkmii on this massive @makermodsai physical AI hackathon Check out @MidcenturyAI at Midcentury.xyz
6
151
Anthropic's study: older, educated, higher earners most at risk. The displacement pattern isn't random. AI replaces application-layer work first. The physical layer operating in the real world is harder. The world is messy and complex and will take time to close this gap.
🚨BREAKING: Anthropic just published a study mapping exactly which jobs its own AI is replacing right now. The workers most at risk are not who anyone expected. They are older. They are more educated. They earn 47% more than average. And they are nearly four times more likely to hold a graduate degree than the workers AI is not touching. The argument is straightforward. Anthropic built a new metric called "observed exposure." Not what AI could theoretically do. What it is actually doing right now in professional settings, measured against millions of real Claude conversations from enterprise users. For computer and math workers, AI is theoretically capable of handling 94% of their tasks. It is currently handling 33% of them. For office and administrative roles, theoretical capability is 90%. Current observed usage is 40%. The gap between what AI can do and what it is already doing is enormous. The researchers are explicit about what comes next. As capabilities improve and adoption deepens, the red area grows to fill the blue. The demographic finding is what makes the paper uncomfortable. The most AI-exposed workers earn 47% more on average than the least exposed group. They are more likely to be female. They are more likely to be college educated. This is not a story about warehouse workers or truck drivers. It is a story about lawyers, financial analysts, market researchers, and software developers. The exact group whose education was supposed to insulate them. Computer programmers showed the highest observed AI exposure at 74.5%. Customer service representatives at 70.1%. Data entry keyers at 67.1%. Medical record specialists at 66.7%. Market research analysts and marketing specialists at 64.8%. These are not predictions. These are measurements of work that is already happening on AI platforms right now. Then there is the pipeline finding nobody is talking about loudly enough. Anthropic's researchers found a 14% decline in the job-finding rate for workers aged 22 to 25 in highly exposed occupations since ChatGPT launched. No comparable effect for workers over 25. Entry-level roles were never just jobs. They were the training ground where junior analysts became senior analysts, where junior lawyers learned how arguments hold together. If that layer disappears, nobody has answered the question of where the next generation of senior professionals comes from. The detail buried in the paper that most coverage missed: 30% of American workers have zero AI exposure at all. Cooks. Mechanics. Bartenders. Dishwashers. The technology reshaping professional careers is completely irrelevant to roughly a third of the workforce. The divide is no longer between high skill and low skill. It is between presence and absence. The company publishing this study is the same company selling the AI doing the replacing. Anthropic had every commercial incentive to soften these findings. They published them anyway. If you spent four years and $200,000 on a degree to land a white collar career, the company that builds Claude just confirmed your job is more exposed than the bartender pouring drinks at your graduation party. Source: Anthropic, "Labor market impacts of AI: A new measure and early evidence" PDF: anthropic.com/research/labor…
2
3
280
Factories are the smartest first stop for robots. They run 24/7 and knock out tedious tasks at peak efficiency. This real factory run only needed 20 min of data for 99%+ success. The real question is how quickly they can take on other tasks and which ones will take more time.
Robots that actually WORK in real factories. Without endless retraining. > Just 20 minutes of data. > 99.4% success rate. > 108 motors soldered in 5+ hours straight. Sub-0.6 mm precision on messy, deformable (!) cables. This hybrid “learning-augmented” system adds neural brains + 3D safety monitoring to ordinary cobots… and suddenly complex factory tasks become reliable, human-safe, and stupidly fast. This is for all of you who are into manufacturing. Huge credit to robotics engineer Yunho Kim @awesomericky99 📍Paper: arxiv.org/abs/2604.22235
3
168
Retention dropped. Expected. The real question is what replaces it. Calculators made arithmetic fall, which freed up bandwidth for higher-order math. If AI handles recall, is the bet synthesis and judgment? Worth tracking.
Students who used AI to study remembered less than those who did not.
1
109
There's a third question nobody asks: how good at what? The ceiling for AI in a text box is different from the ceiling for AI operating in the physical world. Pace of digital AI progress doesn't tell you much about the physical deployment curve. Two different races.
Every AI discussion ultimately rests on two questions: how good can AI get? And how fast? They are predictions about the s-curve shape. Everything else (job impact, potential risks, etc.) is downstream of those questions. I think it would be useful to focus on them more often.
2
93
Good highlights from a qualified and technical room. The data bottleneck is real. Building real-world data infra is the play. But collecting edge cases at scale isn't enough. Full-stack feedback infrastructure is the moat
3
107
Agents bootstrapping their own improvements is interesting, but it points to the real problem. Agents stay locked because they don't get learning loops. Prompt engineering freezes capability. The unlock is agents that gather data from failures and retrain. Not just from better prompts, but from ground truth.
Nous Research built an AI that rewrites its own brain for $2. Most AI agents run the same prompts forever. One open-source repo changes that. Hermes Agent Self-Evolution lets agents rewrite their own prompts, skills, and code. No manual tuning needed. The engine powering it is called GEPA. It's a genetic evolution algorithm for prompts. It uses 35x less data than reinforcement learning. Yet it scores 20 percentage points higher. Here's how it works: 1. The system reads an agent's past task logs. 2. It finds where things broke and why. 3. Then it proposes fixes automatically. Those fixes get tested, ranked, and the best one wins. One full optimization run costs $2-10 via API calls. No GPU required. Safety stays tight: > All variants pass full test suites > File sizes are strictly capped > Changes can't drift from original intent > Every update needs human PR review Prompt engineering is becoming the prompt's job.
3
137
LLM analogy breaks for robotics, real-world experience doesn't transfer as directly. Robotic data is task-specific and costly. No scalable method yet to collect quality data for world models. Industry figuring it out. Capital flows. Infra doesn't. That gap is the game.
𝐁𝐞𝐬𝐬𝐞𝐦𝐞𝐫 𝐏𝐫𝐞𝐝𝐢𝐜𝐭𝐬: 𝐑𝐨𝐛𝐨𝐭𝐢𝐜𝐬 𝐚𝐧𝐝 𝐩𝐡𝐲𝐬𝐢𝐜𝐚𝐥 𝐀𝐈 🤖 1. We're in the GPT-2.5 moment for robotics. Capabilities are real, but the gap between lab performance and field deployment remains wide. 2. Scaling laws are emerging. Data is expensive, capital is the moat. World models may be the shortcut. 3. Talent concentration will crown winners quickly. This is not a market where 50 companies win. 4. Near-term value will accrue to full-stack, vertically integrated players, not pure-play foundation model companies. 5. Defense robotics will produce the first $50B+ IPOs in the category. 6. There will be no robotics bubble. In fact, not enough capital is flowing into the industry. Dive in 🦾 bvp.com/atlas/bessemer-predi…
1
1
119