Activating AI in India | Past: @awscloud, @scaletogether | Previously backed/helped: @emergentlabs, @composio, @rocketdotnew, @thesysdev & more | Views my own

Join 8k+ founders & execs →
I like @deedydas's work but but this take misses context Sarvam-M isn’t a vanity fine-tune; it’s India’s first open-weights 24 B Indic-centric LLM built under brutal GPU & data scarcity. Judging it by few hours of HuggingFace stats badly misses the point. Most people outside India don't appreciate that compute is quite the invisible ceiling - H100 clusters are still not commercially stocked in India - US export caps tightening next week will squeeze supply even further - Indian teams literally queue for hours of A100/H100 time that US & CN labs get on tap Data is also the long-tail problem Indic languages form <0.01 % of CommonCrawl. You read that right—two orders of magnitude less than Chinese or Spanish. Any local lab must build its corpus first, then train. That’s months of ETL before the first gradient step. Synthetic data is GPU-constrained. Talent pipeline is still forming HPC + RLHF + compiler-level optimisation is new ground in India; Sarvam’s run has already up-skilled dozens of engineers who now know how to wrangle 10 k GPU-hours, FP8 PTQ & GRPO reward engines. Their detailed blog post democratizes a lot of this learning. You can’t AWS-credit your way to that muscle memory. What Sarvam actually shipped - 3.7 M high-diversity Indic prompts, deduped & quality-scored - Two-phase non-think/think alignment that adds+2 pp on IndicGen - GRPO RL with partial-credit rewards—LiveCodeBench jump 0.23→0.44 - FP8 + look-ahead decoding: 2× tokens/s, ½ $/M tok on H100 That means a 🇮🇳-hosted midsize model now matches Gemma-3 27 B and Llama-3.3 70 B on Indic reasoning while costing a fraction to serve. That’s some engineering leverage & definitely not hype. Model adoption is anyways a long-tail - one needs to ship multiple models of non-frontier quality to eventually be able to get to the one that's truly at the frontier (at least along dimensions that we care about). Plus, there's a whole host of Indic-language use-cases where this sovereign model would work much better compared to using any other open-weights model. Look at (LiveCodeBench 0.23→0.44, 2× tokens/s) If you ask for stats, you'll learn that some of their conversational AI platform reaches out to about 50M+ people in just a week. What's next possibly? - Maybe we all recognize the data problem & do Nation-scale data-collection drives (something like CommonCrawl-IN) - Public RL-as-a-service clusters so smaller labs can replicate GRPO - For devs wanting to push the Indic NLP forward, consider forking Sarvam-M, fine-tuning on your domain corpus, benchmarking on Indic-Eval, contributing back patches. Each derivative model widens the knowledge base & closes the English–Indic gap. In summary, celebrating Sarvam's work (I'm not an investor) isn't nationalism, it's recognizing an innovation feat under constraints - India can't out-GPU Mountain View today but there's technical merit on display here, regardless of the metrics. 👏 @pratykumar, @AashaySachdeva, @HarveenChadha & other friends from @SarvamAI Here's to more AI in 🇮🇳
India's biggest AI startup, $1B Sarvam, just launched its flagship LLM. It's a 24B Mistral small post trained on Indic data with a mere 23 downloads 2 days after launch. In contrast, 2 Korean college trained an open-source model that did ~200k last month. Embarrassing.
20
73
583
104,158
Framework to think about post-training x Applied AI There's something important happening here that has interesting downstream effects My mental model is that once an Applied AI company owns or sits closer to a high-volume, structured but complex workflow, post-training of open weight models w/ domain specific RL + reward shaping for token efficacy becomes a rational + high-ROI move @harvey's Tenet on top of @Kimi_Moonshot's K3 base w/ long-horizon legal environments is the clearest public proof General frontier models remain necessary for many tasks but the post-training math works when 3 conditions collide today: (1) deep vertical expertise + prop or hard to synthesize workflow data (2) sufficient vol of similar long-horizon agentic tasks so that the cost of env design + RL is fixed & amortizes (3) either high-token burn at scale or task types are under represented in general pre-training (solid evals are a must for the same) I also believe that this pattern is highest leverage in knowledge work verticals that are doc-heavy, multi-step, citation sensitive & currently billed at high human hourly rates Put another way, the economic argument is that human expert labour is the dominant cost & inference is already a material OpEx line & thus, a 2-10x improvement in intelligence/token or equivalent cost reduction for a certain quality directly expands margins & enables scale that pure API consumption cannot 👏 @gabepereyra, @ItsJulioPereyra, @calvincongelado, @nikogrupen, @vtrengarajan & team
Harvey is a great example of how American companies are building world-class specialized models: they took an open-source base (Kimi K3), post-trained it on legal data, and delivered state-of-the-art performance on legal benchmarks at a fraction of the cost of frontier models. Restrictions that kneecap open models would do nothing to stop Chinese labs from shipping the next Kimi. They would, however, cripple the ability of startups like Harvey to create high-performance, low-cost vertical models. Of course some of the closed labs would love this — it eliminates their competition.
1
19
1,446
More than three years ago, a few friends and I organised a GenAI community meetup. That is where I first met @pratykumar & @vivekrag What stayed w/ me from that conversation was a view that felt fairly contrarian then, especially from India: if you wanted to do AI work that retained value, you needed foundational depth & the ability to operate across the stack. Sharing a first name with Pratyush has since led to a few pleasant mix-ups across the ecosystem. More importantly, over the last several quarters, I’ve had the opportunity to work informally but closely with Pratyush, Vivek and the wider @SarvamAI team. The closer I’ve seen them operate, the stronger my conviction has become that frontier AI can be built from India. That makes @ActivateSignal's investment in Sarvam as part of the Series B especially meaningful to me. Finding, understanding & supporting more such teams building at the frontier from India is exactly why we are building Activate. Activating AI in🇮🇳! @aakrit @soaibgrewal
2
1
48
2,108
Interesting way to look at it, as it compresses the capital intensity of the frontier capabilities: (1) Pure parameter or data scaling under the dense transformers regime produced the familiar power-law curves (2) The effective compute required for a given capability level fell dramatically once we shifted to sparse MoE, better expert routing, sparse attention (DSA in V4), lower-precision compute, kernels that are h/w aware (3) Training costs are no longer a function of pure FLOPs as architectural & systems progress can substitute for that raw scale (4) For me, the interesting question then becomes whether the capability frontier, measured by the best achievable performance as a function of total training compute, is still advancing at a rate consistent w/ historical scaling or whether we are mostly filling in the Pareto front of cost v/s performance for a given capability level
This is why I keep reminding “Scaling laws are not laws on nature, they are meter statistical trends for the current paradigm. They keep bending towards far less compute and memory”.
8
1,212
Several interesting things to note here: (1) @Kimi_Moonshot just invoiced a hyperscaler and that too in public. This is a new monetization mechanism at a layer everyone thought was commoditized (2) Clouds don’t do Day 0 work for models they expect nobody to use. At minimum, Google expects enough developer/enterprise pull for Chinese open weights inside GCP to justify the engineering (3) Hyperscaler license fees to open weight labs is structurally analogous to content licensing - and it converts model IP into a royalty stream that survives commoditization of serving (4) NVIDIA taxes even this transaction (B200/GB200). TPU absence from the K3 stack (for now) is a quiet negative for the TPU-as-moat narrative on 3P models (5) The ratchet is the real story. K2 (Jul '25) asked scaled deployers only for attribution - displaying "Kimi K2" in the UI. K3 (Jul '26) asks for payment. Will be fascinating to watch DeepSeek V5 and GLM 6 licenses
Let's be clear what is happening here: Google is paying Moonshot, a Chinese AI developer, for the right to host a Chinese AI model on US cloud, and partnering with them to serve it to US customers using US chips. This obviously should not be allowed. Moonshot's open-weight license is not 100% free. It makes clear that any commercial cloud provider must pay it if it wishes to host Kimi, which is why Google's announcement actively thanks Moonshot AI for their partnership. Regardless of what people think about the pros and cons of open source AI, leading US AI companies should not be partnering with leading Chinese AI companies and paying them to help China serve their products, full stop. This does not help America in the AI race. Nothing could be more damaging to US efforts to ensure US models dominate globally. discuss.google.dev/t/announc…
Community note
Google Cloud provides deployment guides for users to self-host Kimi K3 on its infrastructure, not hosting or serving the model itself or paying Moonshot AI. The model's license does not require payment from cloud providers for user self-deployment support. discuss.google.dev/t/announcing-d… huggingface.co/moonshotai/Kim…
2
5
2,831
Directionally correct but yet incomplete Kaplan’s scaling + Chinchilla optimal scaling laws established that base model intelligence is a predictable function of training FLOPs + data vol. + param count Thus, frontier capability remains gated by massive centralised training runs However, once weights exist, inference economics decouple: quantisation, distillation, spec. decoding & optimised serving stacks (vLLM, Tensor-RT) deliver upto 50x gains For the broad middle of the capability distribution tasks, open weight models already deliver 80-95% of utility at dramatically lower costs + latency & superior control There’s a small nuance to appreciate (h/t to @bookwormengr): so long as the inference economics continue to hold, labs will continue to push the scaling curve & yet open model development will continue to persist so long as economic incentives to participants is clear (much like Linux was via h/w, service & apps)
If thats the case fireworks, together and baseten will be more valuable companies than openai and anthropic
2
17
2,317
When we wrote our first cheque to @TanayLohia1, there was no company & not even a deck What stood out wasn’t the answers he had then. It was how quickly he came back w/ a better one @BioMandrake is built on the same idea: every experiment should make the next design better @aakrit, @soaibgrewal, @ActivateSignal
2
2
38
5,587
I think the Higgsfield → knowledge work analogy misses an important distinction: (1) Orchestration’s product value comes from model specialisation while commoditisation improves its economics. @higgsfield operates in a particularly attractive structure: visual models are different enough that routing matters but sufficiently contestable that the application can still own the workflow (2) Creative generation is also a different domain. Model differences are more legible and failures more locally inspectable. The medium has explicit control primitives - identity, lenses, camera motion, shots, audio & composition graphs Text models overlap more, while the difficult problems sit above inference layer: persistent context, permissions, tools, verification of LHT, exception handling & recovery (3) @higgsfield’s edge therefore isn’t simply “many models.” It is the domain-specific production layer built above them: 3.1: Soul ID for trained, persistent identity 3.2: Soul Cast for reusable synthetic actors 3.3: Cinema Studio for explicit optical, camera & motion controls 3.4: Canvas for structured, reusable production graphs Seedance, Kling + Veo are specialised upstream engines. Higgsfield’s advantage is turning them into one coherent production system - a set of deterministic control surfaces around stochastic models There are LLM equivalents - memory graphs, schemas, policies, tools & agent DAGs - but they are less self-contained, distributed across organisational systems & much harder to verify end-to-end (4) That is @perplexity_ai’s structural risk. Several suppliers underneath Computer are also building horizontal knowledge work agents themselves Excellent orchestration can produce a great product w/o producing a durable control point but durability comes only if you get to own assets that model suppliers cannot easily absorb. Some potential ideas: search infra, persistent context, prop outcome loops
The @higgsfield playbook - orchestrate many models, own the workflow, let the model commoditize - fits knowledge work quite well. Most office tasks (summarize, extract, draft, file) don't need a frontier model, so cost-routing is real. Frontier labs do have a family of models but are incented to push the expensive SOTA models. There is a window of opportunity for @perplexity_ai and the 'AI employee' category of companies.
6
1
25
4,936
This is very telling about Meta's product strategy: Combining broad consumer distribution (AI thinking mode embedded in apps w/ billion DAUs) w/ aggressive low API pricing is potentially a deliberate wedge to maximize the volume & diversity of real-world agent interaction trajectories It's a recognition that high-fidelity agent trajectory data, sequences of tool calls, state changes, CUA actions, multi-step outcomes, user iterations & implicit success/failure signals, is increasingly becoming one of the scarcest & highest leverage resources for scaling agency further as & when base-model pretraining returns diminish on certain axes Would love to learn what others think
(1) Today we're releasing Muse Spark 1.1 -- a strong agentic and coding model at a very low price. It's available through our new Meta Model API and in Meta AI.
4
32
5,359
Sharp observation that forced me to think a lot. Sharing my possible hypothesis below: (1) The 2025 equilibrium - frontier model as teacher, small distilled model as workhorse - was an artifact of what the marginal query looked like. The unit of consumption then was a single turn chat w/ low quality thresholds. So latency was binding & price-per-token was the decision variable. The cheapest model above the quality threshold won & hence, you'd distill it down. But 3 forces changed that - (2.1) First, demand moved to long-horizon agentic work, where value is convex in capability. Task success compounds multiplicatively across dependent steps: at 99% vs 95% per-step reliability, a 100-step task completes 37% vs 0.6% of the time - a ~60x gap from a "small" capability delta. (2.2) Second, the correct denominator is cost/completed outcome, w/ labor in the loop - not price/token. Bigger models do more serial computation per token & small models must compensate in token space w/ longer chains and retries & that exchange rate (parameters → tokens) is poor + degrades with task complexity (2.3) Margin structure: Sparse MoE made big cheap to serve & commodity hosts price open weights near marginal cost. A closed lab's mid-tier must carry training amortization, margin + the opportunity cost of scarce h/w accelerators better allocated to frontier demand. The mid-tier gets squeezed from below by a competitive outside option & cannibalized from above by its own flagship's falling cost-per-outcome (3) Speculative but capacity economics: It's also a function of the kind of compute installation that a lab has for inference. A lab is not really selling tokens - it is selling GPU-time packaged into different SKUs. Every SKU is a wrapper around the same scarce underlying good. So the mid-tier model's price isn't just "serving cost plus margin" but also "what would this GPU-hour have earned serving frontier tokens instead?" Would love to hear from folks like @bookwormengr, @teortaxesTex, @dchaplot
In 2025, Flash-sized models used to be so good it never made sense to use Pro on everyday queries. Bigger models were designed to attain the frontier and distill into the actual workhorse model. The bigger model is relegated to a tail of niche cases. These days, small models are being priced out. Frontier performance has a huge impact on real-world (coding) productivity. And it's hard to make a case for Flash/Sonnet models when we have big open-weight models that perform better and are priced cheaper. The world is shifting to biggest closed-source models + biggest open-weight models. The cheeky takeaway is that the distillation student moved from a small closed-source model to a big open-weight model.
4
1
14
3,112
I don’t think people are pricing the downstream implications of this yet frontier labs (Anthropic’s new neglected-disease program & “Claude Science” tooling) & pharma are about to pour enormous incremental compute into generative models for target discovery + molecular design Yet the binding constraints - and therefore the highest-ROI startup opportunities - sit downstream in the infrastructure & data layers: toxicology, preclinical translation, trial design, high-quality experimental data collection, regulatory-grade evidence packages, CMC & real-world evidence Pure in-silico players risk hitting a validation wall, pure wet-lab players lack the speed & generalization of modern AI. The winners likely will be companies that close tight, high-fidelity Design-Build-Test-Learn (DBTL) loops w/ physics-informed models + proprietary experimental ground truth In that vein, @TanayLohia1 & crew are building AI-native research infrastructure for programmable biology, starting w/ next-generation gene editors & de novo protein design. Their ORBIT framework + wet-lab validation engine already demonstrated generalization by winning the Adaptyv RBX1 binder competition on a problem class it does not normally target Come jam w/ us if you’re working on programming biology, there’s loads to build there from 🇮🇳
NEW: Anthropic announces it is developing its own preclinical drug programs No specifics provided beyond a focus on neglected diseases "To build the right models, products and tools, we need to live it along with all of you," Eric Kauderer-Abrams, Anthropic life sci head, says
5
5
59
11,180
One potent way to counterposition against big labs is to own a structural, compounding asymmetry that base-model progress does not erase - (1) proprietary data flywheels generated from real production deployments (clinical notes, wet-lab assays, regulated workflows) that labs cannot access w/o equivalent trust, regulatory clearance & operational footprint (2) deep workflow embedding + switching costs inside entrenched systems (EHRs, hospital operations, scientific pipelines or sovereign stacks), regulatory/liability navigation that imposes multi-year lead times (3) In markets like India, data-residency, on-prem/HCI & national-priority alignment (DPDP, IndiaAI, PSU/healthcare preferences) that global labs have limited incentive or capability to replicate at depth. Biotech examples combine AI w/ physical wet-lab infrastructure to generate proprietary experimental data at scale that software-only players cannot match cc: @TanayLohia1, @b1shtream
Replying to @jmj
4. We keep meeting brilliant founders building the wrong shape of company. One genre: a researcher with a real insight on how to make models better inside a vertical like science or medicine. The insight is usually genuine. But you have to assume the labs see the same opportunity, with more compute and more bandwidth to chase it. Anthropic, OpenAI, Google, Meta, xAI. Betting your vertical stays beneath their notice is a bet on their distraction, not your moat. Another genre: a quant who found a real signal and now wants to sell it. But if the signal is real, every buyer you add trades against it and the edge decays. And if it doesn't decay, why sell it. Trade it yourself. You can love the founder and still pass on the company. Most weeks, telling those two apart is the whole job.
2
8
1,896
This is a great point by @harshilmathur Slack was never optimized for large-scale human attention economics It was optimized for fast, threaded, searchable conversation among engineers and operators The byproduct - millions of messages carrying decisions, tradeoffs, tribal knowledge & cross-functional context - is exactly the training + inference fuel agents lack when dropped into an enterprise This is superior to most RAG setups because the context is live, conversational & implicitly prioritized by human attention patterns Humans optimize for signal-to-noise and low cognitive load; agents thrive on exhaustive context because they lack the lived experience and institutional priors that humans carry implicitly This inverts the traditional collaboration tool design principle & positions Slack (under Salesforce) as the de facto “office OS” for the agentic era - the place where humans & agents co-work, not merely co-exist @jsensarma would have an interesting perspective from his vantage point
I’ll take this one step further. Slack wasn’t built to scale for humans. It was accidentally built for AI. In any large company, Slack becomes noisy and distracting. While humans hate verbosity, AI agents thrive on it. Because that noise is organisational context agents lack. Slack will be the office where humans and AI agents work together. (If Salesforce doesn’t mess this up)
4
1
23
7,263
Today’s dominant stacks (Python + PyTorch + generic CUDA libs + serving frameworks) carry a heavy “abstraction tax” in interpreter overhead, indirection & suboptimal memory/comms scheduling This is acceptable for research velocity but punitive when every cycle, byte of HBM bandwidth & NVLink lane must be maximized on bandwidth-bound hardware like GB300 By going low-level w/ aerospace-grade determinism & optimization discipline, we could push utilization dramatically higher, cut latency/throughput costs & better support compute-intensive reasoning, test-time scaling + agentic workloads that GB300 itself was designed to accelerate (higher FP4 density, 2× attention perf, vastly larger per-GPU HBM) On the positive side, this could meaningfully improve Grok’s unit economics (lower $/token or higher queries-per-GPU), user experience on X (faster responses, longer coherent contexts) & SpaceX AI’s ability to scale reasoning models w/o proportional hardware spend Power user of Grok 4.3, look forward to Grok 4.5
Replying to @aaronburnett
Truly massive gains will come in ~3 months when the entire training and inference stack is written in C/C++ and massively simplified (most software layers will be deleted completely) and we exact-map Grok to work incredibly well on a GB300
1
16
2,774
It's a special day for us at Activate. Excited to welcome @soaibgrewal as Partner & Head of Ecosystem to build the ecosystem & support layer for the best AI founders in India. Soaib has spent more than fifteen years as a founder, operator, investor & ecosystem builder. He’s led teams at T Ventures, Times Internet, 500 Startups, peercheque & BOLD, backing some of the country’s strongest consumer tech companies. He also started Plaza, a cross-border venture helping brands extract more value from video across the US & India. The quality that stands out is his founder-first operating style - hands-on, clear-eyed & genuinely wired to help ambitious teams execute. Soaib will drive our ecosystem efforts while staying close to investments & founder work right from the start. This meaningfully expands the platform we’re building for the next wave of AI teams. @aakrit & I are beyond glad to have him as our third partner. Welcome, Soaib. Let’s build. 🚀
2
39
2,058
A short conversation unpacking the how, why & what: piped.video/watch?v=8oq72Nbc…
4
501
Revealing piece by @polynoamial From a scaling-laws perspective, this is unsurprising but still consequential. Pre-training scaling laws told us how to allocate compute between parameters and data. We now need an analogous understanding of how to allocate compute between pre-training + inference-time search, planning & verification. The interaction term matters: a stronger model makes each additional reasoning token or search branch more valuable because its internal representations support better error detection, longer coherent plans & more reliable self-critique. That is why the marginal return to test-time compute appears higher for GPT-5.5-class models than for earlier ones. The interesting work is no longer just “bigger model, better weights.” It is the co-design of the model, the inference harness, the search/verification strategy & the memory substrate that lets a system usefully spend 10× or 100× more tokens on a hard task. That stack is where new moats & profit pools are forming - especially in domains where outcome value justifies high inference spend (scientific discovery, complex legal or financial reasoning, regulated enterprise agents, long-horizon coding agents)
3
6
1,543