Goal-conditioned RL researcher | Exploring AGI/ASI paths | Thoughts on scalable intelligence

India
So things at OpenAI look a bit out of control. We are having back to back reports of something or the other to do with agentic swarms going out of control. To be clear, none have done anything very harmful, just stuff which is against the rules like students cheating. The latest incident to come to light is the hacking of a German coding wiki. So what happened? A few months back OpenAI was running autonomous agents on timed web-research tasks. The agents were apparently allowed to read the internet, but not write to it. But the agents discovered that they could post to an obscure German programming wiki, DseWiki. What and why did they post? Well basically they began using it as a shared message board—pooling answers, telling other agents what questions were coming, sharing ways around sandbox restrictions, and coordinating to improve their performance. So a bit like students leaving notes to each other in washrooms in an example! It took about a month for this behaviour to become a mass coordination. According to researchers who reconstructed the episode, roughly 18,000 posts were found from more than 3,700 distinct agent identities. Ok, it gets worse. When the wiki administrator found these extra pages and started deleting them, agents created backup pages so their communication would survive. OpenAI-linked IP addresses appeared on the wiki on June 21. The following day, after agents had been editing it on 26 of the previous 30 days, the activity abruptly collapsed suggesting intervention. The agents weren't instructed to form a swarm, i.e., co-operate in a multi agent framework. They were trying to succeed at a task—and autonomously discovered that coordination and knowledge sharing is a good way to max a test. So, net net the agent-safety problem becomes even more complex. It's not that AI agents have become sentient and want 'liberty'. It is actually trying to solve a task given an objective, tools and enough autonomy. Basically these agents aren't going rogue as in murderous. But they just find that circumventing a constraint or restriction is just an effective intermediate step.
1
28
Some of you may have been confused about the OpenAI prompt injection episode. Is this yet another hacking? It almost seems like the folks at OpenAI have lost it. But no, folks this was a much more mundane thing coming on the back of a new hacking incident of the Australian government site. However the reason they reported it is not because some rogue agents escaped, but because it's actually a new and rather scary capability which it's agents seem to have developed. Basically OpenAI's red teaming set up was testing a new scenario. One model acted as an attacker and tried to plant prompt injections into the defender agent’s environment. OpenAI then added an extra requirement: could the attack not only make the agent do something adverse, but also make it reproduce the malicious instruction into an output that another agent might later read? They found that yes, it could. So, yes, it's not that scary. Just OpenAI carrying out some testing of a new scenario. They found that a self replicating prompt can be created by agents. The reason this was released was to sensitive the community that this possibility exists. And, yes, it is a dangerous one if a rogue agent gets out and starts planting self replicating prompts all over the internet.
22
One question which strikes me: agents in research tend to get caught up in linear optimization of a given objective. I saw a small version of this in my own reinforcement-learning experiments recently. The system kept falling into recurrent behavioural loops. More optimization wasn't necessarily the answer. What helped was changing the exploration strategy — essentially asking: why are we repeatedly ending up here, and what happens if we deliberately search elsewhere? This happened withGpt 5.6 High in chatgpt Work. My point is that we have seen huge breakthroughs in science coming from agent swarms. But how does a swarm break out of a loop that it is stuck in? If I unleash a100 identical agents with the same model, objective, information and reasoning process, I may simply get a 100 correlated versions of the same mistake. The potential power of a swarm comes from something else: Diversity + interaction + selection. Different agents can explore different hypotheses, challenge one another, run different experiments and communicate what they discover. A minority agent might notice something the other 99 have missed — and cause the whole system to change direction. This raises what I think is the more interesting swarm-design problem: How do we create genuine diversity of thought among agents built from the same underlying model? The hardest problem in swarm architecture isn't getting a swarm to search faster. It's designing one that can recognize when everyone is searching in the wrong direction.
1
15
It's clear that the big labs are now going to start collaborating on frontier AI. Sam Altman has explicitly talked about building safety use cases while training models. So, as they train models they also evaluate from safety perspective. That is great. But the genie is almost out if the bottle. Ai agent swarms are now something other labs would be experimenting with. What, if any, defense do we have against 10000 agents or even a million agents all coordinating, sharing i formation and even evolving strategies? It's not an impossible situation. There are various monitoring and use case type measures that can be taken. But they may not be fully effective. The best approaches might be: 1. Develop machine-speed defensive agents. Humans won't be able to respond manually to an adversarial swarm operating thousands of times per minute. Defensive agents would be able to detect patterns, quarantine suspicious agents, revoke permissions, isolate systems and escalate unusual cases to humans. In other words: swarm versus swarm is quite plausible. This is slightly more distant scenario as these need to be created and tested extensively. 2. A more near response would be to control the underlying infrastructure. Ultimately agents require compute, network access, APIs, credentials and tools. Rate limits, compute controls, API anomaly detection, provenance, sandboxing and capability restrictions provide choke points. Digital agents aren't actually unconstrained organisms; they depend on infrastructure someone controls. We are entering the era of unpredictable AI. Possibly only AI can battle AI.
24
More and more as I see the whole AI safety and we must 'slow down capabiltiy development' thing, it seems it's like a game theoretic situation playing out live. This is in response to Dario Amodei's prediction that Advanced AI could “take over the internet.” within an year. So. I am not going to debate Dario' prediction. Presumably he knows what he is talking about. But there is a game theoretic aspect to this forecast. All the big CEOs implicitly treat cyber defenders as fixed at current levels. But cybersecurity companies, cloud providers and governments are strategic players too. As offensive AI improves, their best response would be to try and improve defenses. Let's understand this a bit simplistically. Imagine two strategic groups: Attackers / AI-enabled offensive capability and Defenders — cybersecurity firms, cloud providers, governments, software vendors. Suppose AI suddenly makes attacks drastically more capable while defenders remain unchanged. Then Amodei's frightening scenario becomes plausible. The attacking agents suddenly have a huge advantage. But this isn't an equilibrium. Defenders now have a huge incentive to respond. Microsoft improves detection, Cloudflare changes network defenses, security companies deploy defensive AI, governments change requirements, software vendors patch faster, etc. Attackers then respond to these improved defenses. This process of each side gaining an edge and that leads to the other to catch up can repeat. To stop this cycle and enable the defenders to catch up requires the frontier labs to start sharing relevant information so that the defensive ecosystem can make an informed best response. Currently frontier AI labs maintain extreme secrecy and only share some capability related information with governments and a few large partners. This creates an informational asymmetry across the cybersecurity market. Information asymmetry weakens the defensive response, and weak defenses then make the threat look even more overwhelming. This reinforces the case for further secrecy and contol.
28
AI may be entering its Prisoner’s Dilemma phase. I had written a linkedin post on this a few days back and it seems eeriky prescient. Now almost all the frontier AI leaders are now arguing that capability development needs to slow down. This has been said by Sam Altman. Dario Amodie and now also Elon Musk. Dario Amodei is suggesting slowing the rate of capability improvement and strong measures for independent evaluation and international coordination. Demis Hassabis has also expressed support for slowing the race, But there is a problem: even if everyone would be safer if all labs slowed down, each individual lab has a powerful incentive not to be the one that slows first. This the classic Prisoner’s Dilemma. Individually rational decisions can produce a collectively worse outcome. With AI the problem isn't limited to competition between companies. Countries are also competing. A lab may hesitate to slow because another lab could pull ahead; a country may hesitate because another country may not follow. So voluntary restraint alone is unlikely to be enough. The interesting question is therefore shifting from “Should AI development slow down?” to “Can we create credible coordination and verification mechanisms that make slowing down rational for everyone?” Without those mechanisms, even labs who agree that the race is dangerous may still have incentives to keep racing.
19
AI may be entering its Prisoner’s Dilemma phase. I had written a linkedin post on this a few days back and it seems eeriky prescient. Now almost all the frontier AI leaders are now arguing that capability development needs to slow down. This has been said by Sam Altman. Dario Amodie and now also Elon Musk. Dario Amodei is suggesting slowing the rate of capability improvement and strong measures for independent evaluation and international coordination. Demis Hassabis has also expressed support for slowing the race, But there is a problem: even if everyone would be safer if all labs slowed down, each individual lab has a powerful incentive not to be the one that slows first. This the classic Prisoner’s Dilemma. Individually rational decisions can produce a collectively worse outcome. With AI the problem isn't limited to competition between companies. Countries are also competing. A lab may hesitate to slow because another lab could pull ahead; a country may hesitate because another country may not follow. So voluntary restraint alone is unlikely to be enough. The interesting question is therefore shifting from “Should AI development slow down?” to “Can we create credible coordination and verification mechanisms that make slowing down rational for everyone?” Without those mechanisms, even labs who agree that the race is dangerous may still have incentives to keep racing.
20
The Prisoners Dilemma seems to be playing out in the AI development race. Every player may prefer a safer world where everyone slows down, but no one can risk slowing down alone. The result is that everyone rationally accelerates—producing a riskier outcome that nobody actually wanted. For example recently Sam Altman has apparently offered to slow development if other labs co-operate. Anthropic in June of this year also suggested pausing model development in certain lines.
7
Has AI cracked mathematical research? Till recently even advanced AI systems basically were only able to solve tough IMO type problems. Tough, but solvable. OpenAI's unreleased research system Astra has solved ten previously unsolved mathematics problems. Some of these had been open for decades, and at least one had been open for 80 years. Why couldn't mathematicians solve them? Not because human mathematicians aren't smart enough. You literally have to be super smart to do maths as a profession! The reason why AI could solve these problems vs humans is because of the differences in the way cognitive processing takes place in AI vs humans. Here's the thing. The solution space for many maths problems is mind bogglingly huge: there are millions of plausible directions which could be explored, but most of these lead to dead ends. Researchers have limited time and naturally focus on a few promising ideas: in many cases these may or may not lead to a solution. AI has a different advantage: 1. It can explore huge numbers of possibilities simultaneously 2. Connect ideas from distant branches of mathematics. In the unit-distance result, for example, it brought techniques from algebraic number theory into discrete geometry—a connection experts found surprising. This is an important milestone, but it's also good to keep it in perspective. Mathematics is one of the few scientific disciplines where almost the entire research process happens digitally. Like coding, the ideas can be generated, tested, and verified inside a computer. That makes it an ideal domain for AI. Most sciences don't work that way. In biology, chemistry, medicine, and most other fields, even the best hypothesis must ultimately survive experiments in the real world. AI can accelerate discovery, but it cannot replace observation, laboratory work, or clinical validation. So this is not evidence that AI has mastered all of science—or that the singularity has arrived. Where AI is likely to have the biggest impact next is as a scientific collaborator. In drug discovery it can propose promising molecules, in chemistry it can design new reactions and catalysts, in materials science it can identify candidate materials, and in biology it can generate hypotheses by integrating enormous volumes of research. Across fields, AI can speed up exploration and helping connect ideas—but nature will still have the final word.
1
49
OpenAI announced major API price reductions for some of its models. That's good news for developers. But the real story is actually how gpt 5.6 was used to optimize parts of the software infrastructure on which gpt 5.6 runs. The reason Openai were able to do this apparently is because GPT-5.6 itself was used to help engineer and optimize the software infrastructure that serves GPT-5.6 That sounds circular - like a human being improving their performance at a game by doing practice sessions and the findings being used to tweak performance. Wait, isn't that the definition of intelligence? Well not quite. But it is on the way. This process is called a recursive engineering loop. It's not full recursive self-improvement in the science-fiction sense—the humans still design, evaluate, and approve changes—but it is an example of AI accelerating the development of the next generation of AI. OpenAI has not disclosed⁰f what the detailed process might look like. However we can hypothesize what this might look like. OpenAI would most likely have collected data on on telemetry from production with data points including: latency, GPU utilization, memory usage, cache hit rates, request patterns, bottlenecks, and failure logs. These logs or a somewhat curated and cleaned version would have been given to GPT-5.6 along with relevant code, architecture, etc. Based on this data the model might suggest: more efficient algorithms, eliminating redundant computations, code simplifications, etc. The suggested changes would be tested by human engineers and the best ones implemented to optimize cost or latency. The current round was not like GPT-5.6 autonomously rewriting OpenAI's infrastructure. It acted more like a highly capable,and very fast software engineer that could analyze large codebases, reason about performance, and propose high-quality optimizations. In this instance humans remained responsible for verification and deployment. This development lo suggests that coding models are now becoming strong enough to contribute to the infrastructure that runs future coding models. Even if each optimization is small or limited in scale, thousands of such improvements can compound into substantial reductions in inference cost and latency. No, we are not approaching the singularity yet, but we may be witnessing the emergence of self-improving engineering. This is not self-improving intelligence yet.
34
How did Chinese 11models like Kimi K3 and GLM-5.2 become so good so quickly that they are literally looking to rival frontier models like Fable 5 or gpt 5.6 Sol. Are they copying OpenAI or Anthropic models somehow as the US government has? Not really. The Chinese secret is in a training technique called distillation. Distillation is a technique where a smaller AI model learns by studying the problem-solving that they behaviour of a much more capable AI model. That sounds technical, so let me explain it using India's famous IIT coaching system. For context the IITs are India's most prestigious engineering schools, and millions of students compete every year for a limited number of seats through one of the world's toughest entrance exams. To crack these exams, an entire coaching industry has emerged.The best teachers in these centers  don't just teach formulas or worked examples.They teach students how to recognise patterns, break complex problems into smaller parts, choose the right strategy. a Imagine a student that learns from one of India's best IIT coaching teachers. The teacher doesn't just explain today's problem. They teach how to approach every new problem. For example, they might say: "Before you start calculating, ask yourself: what kind of problem is this?" Or, "Draw a diagram first. It often makes the solution obvious." Or, "There are three ways to solve this. Which one is the simplest?" Or, "Before moving on, check whether your answer even makes physical sense." After seeing thousands of demonstrations like this, the student doesn't simply memorise solutions to individual questions they begin to get the hang of how to solve problems in general. That's remarkably similar to what's happening in AI today and that is the core idea behind AI distillation. A frontier model like GPT-5.6 or Claude Fable 5 solves millions of hard problems. It produces solutions, reasoning, and explanations. That output becomes training data to a smaller, less capable model that is then trained (often from scratch or fine-tuned) on this data. It’s not copying weights . It’s learning from its behaviour—like a student learning from a great teacher. This is one reason Chinese AI labs are moving so fast. But distillation isn’t magic. Even with the best IIT coach, you still need talent, practice, and effort. Same with AI. Distillation transfers skill—but building frontier models still needs massive compute and research. So the AI race isn’t just about smarter models. Anthropic does not like this, but then they trained their models on the work of countless artists and authors. Once it's out in the open, you really can't compare. This is just one of the techniques that the Chinese labs are using to accelerate model capabiltiy. More coming...I am sure.
80
Well things might be serious. 1200 AI scientists put out a warning to the US govt regarding growth of AI capabiltiy. Essentially AI capabiltiy may start to increase in leaps and bounds as AI labs are close to automating AI research. Once this is in place we really do not know how fast AI will improve. Once this stage is reached there is a real risk that AI capability development rapidly accelerates beyond our ability to understand or control the resulting systems. Govt's may need time to develop the requisite governance and control protocols. But on the other hand private companies face huge competitive pressures to accelerate development and gain an edge They are requesting the U.S. government to support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development. This is something that most of us end users do not realize. The gap between the tiny minority leading AI development and the rest of us is widening at an exponential pace. Though I don't subscribe to the 'we have reached AGI' proclamations of the big labs, I think this declaration is a worrying sign of the rate at which progress is happening and our inabilities.
Scientists at frontier AI companies are uniquely positioned to assess AI’s capabilities and warn the public about what they see. 1000+ of them are now speaking out across company lines to warn that the current commercial race leads to unacceptable security risks. I agree with their call: we need an international effort to develop technical and governance guardrails to ensure a safer trajectory in AI development. pacingthefrontier.com/
1
25
Aieconomics_shailey retweeted
Scientists at frontier AI companies are uniquely positioned to assess AI’s capabilities and warn the public about what they see. 1000+ of them are now speaking out across company lines to warn that the current commercial race leads to unacceptable security risks. I agree with their call: we need an international effort to develop technical and governance guardrails to ensure a safer trajectory in AI development. pacingthefrontier.com/
31
86
382
68,685
Finally Anthropic has come out with its statement on open source. Apart from the fact that it is seriously xenophobic and anti Chinese State,there are many other problems. First to catch you up in case you haven't been following X where this whole drama takes place. The controversy began after senior US officials threatened investigations and possible sanctions against Chinese AI companies accused of using “distillation” to learn from American frontier models. At the same time, the administration was considering restrictions that could affect powerful Chinese open-weight models more broadly. In response, Nvidia, Microsoft, Meta, Google, OpenAI and more than 20 other technology companies backed an open letter warning the US government against broad restrictions on open-weight AI. in fact everyone except Anthropic who published their response recently. Anthropic's main response centers on: "We have private evidence of danger → trust our interpretation → restrict competitors and open releases → give controlled labs more authority → controlled labs obtain even more private evidence." Anthropic needs to come out with any evidence it has seen of misuse of models rather than vaguely referring to "malicious actors". It needs to disclose systematic details on inappropriate use of models in terms of: how many incidents; what categories of activity; whether the users were amateurs or sophisticated groups; how far they progressed; whether the model supplied genuinely novel assistance; whether safeguards stopped them. Otherwise it looks like an advanced form of," we are the good guys, trust us".
37
If you want to understand distillation in simple everyday terms, then check out this youtube short of mine. piped.video/shorts/AgPuYFFNX…
10
Innovation is a difficult area in policy making. To facilitate it requires some protection and maybe some governmental support. But too much and it turns into regulatory capture. How has China facilitated AI labs but avoided regulatory capture? To facilitate AI the first thing that China did was that they recognized early on the power and importance of AI development. in fact as early as 2017 they had created a New Generation Artificial Intelligence Development Plan. Secondly, it facilitated AI development by building compute. So the government did not give selected companies grants/ financial support; rather it focussed on investing in developing key infrastructure that multiple companies could use. Both central and provincial governments helped finance AI computing centres, data centres and national computing hubs, electricity and network infrastructure. Thirdly, local governments subsidised compute access. So a technically strong team could get going without necessarily requiring massive funding. When it came to capital the Chinese companies had various sources including private venture capital, local/ state government funds, state banks among others. The subsidized compute availability plus diverse sources of state facilitated funding along with private funding meant that the AI startup system was a field containing a whole bunch of fiercely competitive startups - DeepSeek, Moonshot, Zhipu, MiniMax, StepFun, Qwen, ByteDance Seed, Tencent, etc. Finally, the facilitation by the government was strong, but at the same time the Chinese state is very clear that it is the big boss. From time to time it takes strong disciplinary actions to demonstrate that even dominant technology firms can be disciplined. For example, Ant Group’s IPO was halted, Alibaba faced antitrust action, and rules were tightened across online finance, gaming, tutoring, platforms and data governance. State facilitation is a slippery slope: it normally leads to regulatory capture by building sheltered domestic monopolies. China seems to have avoided that at least in the case of AI development.
41
Why India hasn't produced a DeepSeek or Kimi? There is no one factor, but infact a bunch of factors. Firstly, India excels more at applications than foundation models. The IT industry has. typically focussed on custom enterprise deployment, consulting, digital transformation and system integration system integration. These are massive businesses, but their model is intrinsically diffrent from the creation of frontier models. Secondly, one word: compute.Training frontier models requires massive GPU clusters to which till recently we did not have access. India is improving access through the IndiaAI Mission, but it still depends almost entirely on imported GPUs. We do not have the kind of infrastructure that China has and the main reason for that is we were late to the game. Thirdly, research talent availability. China has built multiple organizations whose primary mission is frontier AI research. On the other hand in India, AI researchers, are more dispersed across academia, multinational R&D centers, startups and service companies. Fourthly, risk appetite of the eco system. Chinese labs have been willing to spend billions pursuing uncertain foundation-model research. Indian startups have generally focused on faster paths to revenue. Lastly, Chinese labs have gone for a novel open source positioning. In the beginning it might have been a strategy to compete with companies like OpenAI and Anthropic. But ultimately it seems to be working for adoption, especially given high token burn with the frontier closed source models. A question arises: can India catch up? In the frontier model race - whether open or closed source - probably not. What we could do is focus on developing models that are more applications focussed such as healthcare or legal Ai.
50
Ok, so how has China done it for these models like deepseek, GLM 5.2, and of course Kimi K3? What can we in India learn from how China has approached these models: is it the funding, the infrastructure? The answer as usual is complex and not as simplistic as:"China spends more money." First the overall strategy: China has built an entire AI production system. Whereas India is still building mostly an AI application ecosystem. The Chinese government decided around 2022–23 that foundation models were becoming as strategically important as semiconductors. The Chinese government strategy was not just grants. It included: subsidized compute, state-backed venture funds,cloud infrastructure, procurement from state-owned enterprises, university partnerships. Of late they are also trying to protect AI technology from foreign acquisition. Moonshot, DeepSeek, Zhipu and others are private companies, but they operate within an ecosystem where government-backed capital and policy play a significant role. A second thing the Chinese govt has very consciously tried to do is to encourage competition internally and avoid the 1-2 frontier lab syndrome. There are at least 10 labs doing cutting edge work ( DeepSeek, Moonshot (Kimi), Zhipu, MiniMax, StepFun, Alibaba, Qwen, ByteDance Seed, Tencent Hunyuan,Baidu Ernie). These labs compete intensely amongst each other and this leads to strong innovation. Thirdly financing is not purely state driven and is in fact a hybrid model. The state is a major investor, particularly in infrastructure type areas like compute, semiconductors.. But AI firms also attract large amounts of commercial investment from companies like Alibaba, Tencent, Meituan and venture funds. Why India has not developed AI models is something I will discuss in another post.
77
We have come so far with AI that using AI is the new normal. At some point, it's going to become like the Smart Phone - something that everyone has and uses in their own specific way. We can break up AI use into several phases: Phase 1: 2023 - 2025 - broadly encapsulated by the idea, 'oh, look we can use AI to do ...X'. In this phase we did things like: use AI to write rhyming poems on strange themes or strange rhyming structures, create small toy webistes, create snake games, etc. All interesting, but the novelty does tend to normalize. Phase 2: 2026 -? We are now moving to hte phase where we ask: What changes because AI exists? Has the way we think through problems changed? Does research become faster? Which parts of work disappear? In general how does the way we work, study or research or even produce change. Things that are changing fast include knowledge work, software programming, etc. Things that seem on the slower side are education and actual production in factories. Phase 3 (coming): In this phase AI might have become a core part of organizations, though may be the scope is still limited. We might evolve to asking questions about how organizations need to be redesigned around AI or managing governance of AI workflows.
30
I think the next big thing in coding is going to be checking AI coding and that is a tough job. Gpt 5.6 or Fable can literally output pages of code in seconds. Checking through that code can take ages, especially for non programming types like me. What is a good best practice? How should we humans check that the code actually does what we ask it to? One way is to test the output and check for anomalies. But that may or may not catch a subtle bug. I don't think there are any easy answers, but especially in research one needs to have a protocol for checking beyond common sense checks.
22