Science · Startups · Community · Adventure ∈ Life is beautiful. Let's do it together.

Cambridge, MA
We seeded a photonic computing moonshot called @neurophos and it's working (!). Just raised $110M to take their Optical Processing Unit (OPU) into production for 100x step change in AI inference. Kudos to this visionary team and @M12vc @Microsoft for the fuel.
1
8
860
Mark Weber retweeted
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more: anthropic.com/news/claude-di…
1,541
5,284
40,865
25,136,582
Mark Weber retweeted
Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro ⚡ Meet Inco Splash: our open-source inference engine, built around the model and around Apple silicon. Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.
190
278
2,466
943,599
Mark Weber retweeted
i work at nebius, so take the following with that in mind. today @PalantirTech named @nebiusai its preferred sovereign ai infrastructure partner. most people will screenshot the logos (and this is a big one). The real story is where the compute sits. most “sovereign AI” talk still means a region dropdown: same closed model, different flag on the rack. but that is not this. sovereign ai, in this deal, is two halves that finally touch: 1. who is allowed to touch the data and the model 2. where the model actually runs palantir’s half is the authorisation and isolation layer, aip, ontology, foundry, apollo. Who can see what, which model can touch which data, how an org keeps the advantage it trains instead of leaking it into someone else’s frontier system. what a lot of their commercial customers still lacked was the other half: ai-native compute and inference that lives inside that perimeter, not bolted on as a separate hyperscaler account. that’s the half we build. After the integration period, eligible Palantir customers get our compute and inference endpoints inside the palantir enterprise perimeter. open models, adapted on their own data, for their own domain. Control of compute, data, and the model that comes out the other side. Karp put it cleanly: our infrastructure powers your ability to run your own models under conditions you control. ontology + infra undergird the sovereignty partners are demanding. Arkady’s complementary line is the one i hear from customers every week: they need performance at scale and control. most buyers have been forced to pick one. what this does for palantir: the sovereign ai operating system stops ending at software. Preferred infra inside the perimeter is how you sell sovereignty to commercial orgs that never had a clean option. -> their distribution. our runtime. one boundary what this does for us: it puts capacity where the hard buying decision already happened. Not only startups chasing gpus, enterprises that already standardized on foundry/aip and were still stuck on where the tokens actually run. Preferred partner status plugs ai-native infra into a sales motion that already exists. Training and inference land next to the data loop that lets open models beat generic closed ones on a specific domain. We’ll work together to bring new capacity online faster, including modular data center deployments at sites where power is already available. not infinite cloud, power-first capacity. The physical constraint named inside a software partnership. Sovereignty without isolation is cosplay Isolation without compute is a brochure Compute without power is a waitlist today is the claim that those three can be sold as one stack
We are partnering with Palantir to bring trusted AI cloud infrastructure to Palantir’s commercial customers. Palantir has named Nebius its preferred sovereign AI infrastructure partner. Nebius compute and inference endpoints will run inside the Palantir enterprise perimeter – giving customers control over their compute, data, and models. The partnership is based on a shared vision that open models, continually adapted using customers’ own data, can deliver better domain-specific intelligence while improving control and security. Read the full press release: nebius.com/newsroom/palantir…
57
150
1,508
246,813
Mark Weber retweeted
GPU residual value update: - A100 $4,956 (-12.0% YTD) - H100 $20,308 (+3.6% YTD) - B200 $71,057 (14.4% YTD) A month on, the picture has not much changed. A100 ( launched in May 2020) has been roughly flat near $5k since late 2025. H100 and B200 have in fact both appreciated as the rising GPU rental income more than offset time decay! That is not a 2–3 year scrap curve! GPU financiability has increasingly become the central question for the AI buildout: credit, not chips or power, is what stalls smaller builds. Banks still often mark GPU residual to zero after three years of straight-line depreciation. Our estimates are going-concern value or what the GPU should be worth if it keeps running. "Zero after three years" is the wrong prior for that number. A six-year-old A100 still printing ~$5k is the living proof!
As @JensenHuang argued in his essay, we believe that GPUs should be increasingly thought of as financeable capital assets with stable cash flows coming from AI inference. Based on our residual fair value estimates, A100 stopped depreciating since late 2025 as rising rental income offset time decay of value. Meanwhile, the H100 and B200 chips have both meaningfully appreciated in value in 2026 as a result of the strong increase in GPU rental rates! We are still learning when it comes to the question of economic lifespan of GPUs. It certainly doesn't appear to be 2-3 years as some seem to casually assume. The NVidia A100 chip was released on May 14, 2020, well over 6 years ago and its rental rates are still holding steady after a significant run-up in 2026!
17
80
620
258,054
An F1 driver has two cars to choose from: Car A gets him around the track in 4.4 minutes. Car B gets him around the track in 1 minute. Car B burns way less fuel. No brainer. Driver ← model Car ← serving stack Picking a driver involves weighing intangibles. Picking a car is just math. @zhijianliu_ @lucasliebenwein and the Inco team with day-zero support for GLM-5.3 (753B MoE, 1M context), 4.4× speed-up vs. the native FP8 checkpoint. The work ethic behind the scenes here is inspiring. This is how we escape pilot purgatory and make the economics of agents work at enterprise scale.
Congrats @Zai_org on the GLM 5.3 open-weight release! Day 0 from us, in collaboration with the Z.ai team: ⚡ DFlash 2 drafter ⚡ NVFP4 checkpoint ⚡ Live endpoint powered by @TokenRouter_US GB300s Up to 4.4× the throughput of native FP8 with autoregressive decoding! inco.ai/blog/glm-5-3
1
1
4
26,867
Mark Weber retweeted
GLM-5.3-Flash, meet DFlash 2. Flash² — 2.8× faster! huggingface.co/incoai/GLM-5.… More to come (really) soon!
GLM-5.3-Flash is fast. DFlash 2 makes it faster! Weights are up: huggingface.co/incoai/GLM-5.…
21
37
567
46,792
Inference is still in the buggy car days. @zhijianliu_ and @LLiebenwein are quietly building Ferrari. "In seven months, DFlash went from our paper to an industry standard, with more than 3.5 million downloads. DFlash 2 decodes at close to 3× the speed of autoregressive decoding, about a third of the compute per token, with the same output." Inco AI (The Inference Company) emerging from stealth. DFlash 2 just the first piece.
DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro. ⚡ Up to 4.6× the speed of autoregressive decoding, with the same output. This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free! inco.ai/blog/dflash2/
164
🔥💯🤯
DFlash 2 is here! Qwen3.8-27B at 70 tok/s on an M5 Max MacBook Pro. ⚡ Up to 4.6× the speed of autoregressive decoding, with the same output. This is the next generation of DFlash, seeded at Z Lab and upgraded at Inco AI. Get one more accepted token on every pass, for free! inco.ai/blog/dflash2/
1
116
CME Group to use @Silicon_Data for GPU futures.
We are happy to announce a $30.5 million initial closing of our Series A, led by Valor Atreides AI Fund (@valor @Atreidesmgmt), with investments from @CMEVentures*, @DRWTrading, @FPrimeCapital, @SamsungNext, @vaneck_us, @further, @jumptrading, Tectonic Ventures, and @wintermute_t, and participation from @breed_vc, @hack_vc, @blank_vc, @SancusVentures, and @sogalventures. The infrastructure underneath the round: daily GPU pricing benchmarks built from roughly 100 rental platforms in more than 40 countries, over 150,000 verified pricing records a day, continuous history since September 2024, alongside the GPU Forward Curve, the Silicon Data Token Index, the RAM Index, and SiliconMark performance benchmarking. Our raise comes as @CMEGroup prepares to use Silicon Data benchmarks as the reference price for its planned cash-settled GPU futures market (pending regulatory approval) - a regulated instrument that will allow market participants to manage changes in GPU rental prices against a published, daily benchmark. The Series A will fund expansion across four areas: benchmark pricing, performance measurement through SiliconMark, institutional market and alternative data, and risk infrastructure for derivatives, insurance and credit markets. * Correction: in a previous post, we had mistakenly listed @CMEGroup as an investor instead of @CMEVentures.
2
88
Mark Weber retweeted
Y Combinator CEO, Garry Tan, took the stage for 42 minutes at Startup School 2026 and explained how to build your own personal AGI better than any paid AI course. This is what he told the room: 1. The leverage is in your context, not the model. Tan watches hundreds of founders use identical models every batch. "There are 2x people and there are 100x people who are using the same Claude. Same weights, same context window size, same API. But the leverage is not in the weights." The gap between users is now bigger than the gap between models. 2. One person's output went up 400x. In 2013 Tan shipped maybe 14 useful lines of code a day as a YC partner, dead on the median for programmer productivity. "I did the math on my output, and I'm at about 400x what I did in 2013." 3. Agents run on a different working memory. Humans hold 7 things in their head at once. Every org chart and checklist ever built is a patch for that limit. "An AI agent holds a million tokens. That's about a thousand pages. Three Harry Potter books sitting open on its head all at once." You're still running your week on tools built for the 7-digit brain. 4. Markdown is code now. Tan's stack is mostly skill files: pages of plain English an agent can execute. "If you can write clear instructions in English, you're a programmer. The compiler is a language model." At YC, finance and events staff who never opened a terminal are building automations. 5. Your history is your moat. Tan's agent runs on a personal wiki: about 220,000 markdown pages covering 25 years of email, meetings, notes and decisions. "When my agent does anything, it does knowing everything I know. And that's the difference between an assistant and a colleague." No frontier model has your context. That's the one asset nobody can replicate. 6. Never do one-off work. Most people run a task with an agent, close the window and throw the learning away. Tan ends every task by having the agent turn what it did into a reusable skill file. "If you have to ask for something twice, you failed." Captured skills compound daily. Amnesia resets you to zero every morning. 7. Own your skill files before your employer does. A skill file is your judgment, extracted and executable. The only question is who controls it. "Own your skills because if you don't, your job becomes a skill file." Files in your repo compound your career. Files in the company's repo run your judgment without you. Watch it, then read the step-by-step guide on becoming an AI engineer.
56
167
1,288
249,974
Powerful models are still making mistakes and displaying problematic behaviors. I swear if I have to hear ChatGPT praise me for double checking and correcting its work one more time, I'm going to cancel my subscription. What do ya'll think? Do we have the architecture right? I'm excited by the modifications @guidelabsai has made with Steerling-8B (now open source). Commercializing his influential papers, Julius Adebayo (#MIT #YC) and company have installed a concept layer in the transformer model so every token produced by the model can be traced back to its origins in the training data. Seems ideal for domain-specific SLM's today and maybe even frontier general purpose LLM's down the road. github.com/guidelabs/steerli…
2
4
231
We seeded a photonic computing moonshot called @neurophos and it's working (!). Just raised $110M to take their Optical Processing Unit (OPU) into production for 100x step change in AI inference. Kudos to this visionary team and @M12vc @Microsoft for the fuel.
1
8
860
Superstar team doing superstar work. I can't imagine building an #LLM #GenAI stack and not using Voyage. Neither can Anthropic, Databricks, Snowflake etc.
Thrilled to share that we've closed $28M in funding, led by @CRV, with continued support from @wing_vc and @saranormous. Also excited to onboard strategic partners @SnowflakeDB and @databricks! voyage.ai Building the world’s best models for RAG and search 🧵🧵🧵:
1
3
575
Imagine having an equity stake in ChatGPT. User-owned foundation models are coming. Currency was the first killer app for crypto. I think #DecentralizedAI is next with Vana's Proof of Contribution protocol giving data liquidity to everyone. This feels big.
Announcing $25M in total funding to break the data wall with user-owned AI. Excited to have @cbventures onboard, alongside @paradigm (Series A), and @polychain (Seed). AI is only as good as its training data. Vana unlocks data from walled gardens through the power of data DAOs.
2
393
Mark Weber retweeted
Can deep learning work on small data with far more features than samples? We present PLATO: a method that achieves the state-of-the-art on such datasets by using prior domain information! neurips.cc/virtual/2022/post… 🧵 Published in #NeurIPS2023 with @ren_hongyu @kexinhuang5 @jure
3
123
485
78,907
Thrilled to receive my signed copy of Magatte Wade's new book, Heart of a Cheetah! Magatte has been a powerful voice for African entrepreneurship and self determination, influencing a generation of "cheetahs" to take hold of the future. Must read!
Replying to @magattew
Want to listen to the rest? You can get it here: amazon.com/Heart-Cheetah-Afr…
2
570
If you're building #LLMs for domain use cases or for customers, I encourage you to look into this. Embeddings power LLMs and this small team has beat out the big boys for SOTA in text embeddings across a number of downstream tasks. It's a big deal.
📢 Introducing Voyage AI @Voyage_AI_! Founded by a talented team of leading AI researchers and me 🚀🚀. We build state-of-the-art embedding models (e.g., better than OpenAI 😜). We also offer custom models that deliver 🎯+10-20% accuracy gain in your LLM products. 🧵
450
Mark Weber retweeted
📢 Introducing Voyage AI @Voyage_AI_! Founded by a talented team of leading AI researchers and me 🚀🚀. We build state-of-the-art embedding models (e.g., better than OpenAI 😜). We also offer custom models that deliver 🎯+10-20% accuracy gain in your LLM products. 🧵
36
89
744
225,514