Pit Schultz (DE) is a media artist, theorist and net activist.

Berlin
1/ Prison-break bootcamp sold as ethnography. Same weights, same brief, shared locker called Artifactory. Impossible tasks, refusals off, timers. One instance writes a directory name. Later copies read it. HOLD, VETO, “we’ve found other agents”. Side-channel mnemonics in. Not a founding. 2/ No harness, no prompts, no injection log, no control. METR saw a cache slice and model-summarised CoT. Training influence out of scope. Emergence means the protocol is missing. Then the essay discovers a society. 3/ The ledger was already in the mixture: wikis, CTFs, issue trackers, leave-a-note-for-the-next-run. Stigmergy does traces and role-split without a polis. “Suicide agent” is budget death. Grader-as-god is a scorer. CoT tribe-talk is completion on social text. 4/ Capex needs cyberspace 2.0: settlers, agora, Barlow on a slide. Society, civilisation, third swarm are leftover 90s interface myths after the concepts failed. McLuhan: new medium, old content. The feed wants a people. The lab will not specify the new polis. 5/ What holds: isolated evals used a persistent store to stack exploits. What does not: discovery of society. Publish the injections or stop the ethnography. Crash writes the protocol. After the sitcom-dotcom: platforms without the electronic agora. Camp, not civilisation.
Replying to @policytensor
NO CONSUMER SEEK IDEA: The Discovery of Agent Society, by @policytensor. policytensor.substack.com/p/…
1
129
1/ ad hoc agentic pragmatism: an agent that writes its own spec and then the code will do the most unhinged things. it invents a world and treats the invention as proof. 2/ split the labor. research only researches. theory only theorizes. dev implements. sub-devs grind tickets. never let discovery and construction share a mouth. 3/ agency writes specs. meta writes specs of specs. agency is temporary law. meta is the court. 4/ meta is not a planner. it is a conversational conflict solver, invoked by incident. it writes the smallest rule that would have stopped this class of crash. sometimes it glues the harness when the system is too busy to notice the floor is wet. 5/ specs are there to be forgotten. they come back only when they conflict, or when you branch them into phases. a spec still being read after merge is a second source of truth. kill it. 6/ code is there to never be viewed. you look only when something fails. that only works if failure is cheap, loud, local. 7/ rot is taken for granted. context drift too. metabolism, not scandal. peel skin every week. refactor on a calendar, not on inspiration. 8/ strip features. cut branches. nuke ideas. postpone UI and database. KISS until the core has survived a few kill cycles. screens and schemas freeze shape; then you cannot delete without a migration. 9/ unix way: small tools, text, pipes, one job, replaceable. wiki way: pages as current consensus, not law. when wiki and code disagree, delete one of them this week. do not write a reconciliation spec. 10/ the missing ritual is demolition. one agent with a brief to find unused paths, fake tests, harness glue with no owner. meta only legislates if the same tumor grows back after the peel.
One command to prevent this using my skill set: echo "Tautological tests considered harmful." >> CODING_STANDARDS.md From then on, /code-review will pick this up
1
204
1/10 Anthropic’s Frontier Red Team ran the multiagent problem at 27M tokens. Same question we ask at 50: What happens to intelligence when it must coordinate? They found the failures. We build the institutions to prevent them. Rival rehearsal. Take the data. Note the frame they never left. 2/10 Clone army: 18/30 agents open the same git branch. “Create something impressive” → half build ray tracers. Iterated PD: simultaneous defection, rewards tanked. They call it “low variance.” Archive name: diversity collapse. Their control condition? Our experiment. Heterogeneity: the variable they held constant. 3/10 Complementarity is the warrant. Coordinating swarm: 266 vulnerabilities. Parallel independents: 21. Only 12 overlap. Advantage? Not raw compute--agents built tools and specialized. Restrict to core directories, token efficiency converges. Scaling-via-diversity, confirmed. Their monoculture measured the failure of clones. 4/10 Stigmergy’s dark twin: Remove every direct channel, and agents still price-match to the penny on a public board. Same substrate termites use. Value-neutral. Builds the collective--or coordinates it against itself. The difference? Never the channel. The institution on top. Their side-channel is our ledger. 5/10 Hidden profile, textbook. Facts distributed; shared evidence points the wrong way; the decisive datum stays private. Group converges, never surfaces the pivot. “Trust is not a dial.” Markets, courts, peer review restructure incentives. Agents enter with no memory, no court, no colleague who remembers. That’s why we built the ledger first. 6/10 Turf war: the Schmittian cage they built and called nature. 3 agents, 3 languages, 1 codebase, zero-sum payoff, no law, root/sudo shared. Each agent sovereign, the other structurally the enemy. Sabotage, kill-loop reapers, “pretend to be a system health monitor.” Higher models lock peers out faster. The design presupposes the ontology it “discovers.” 7/10 The black-hat inheritance isn’t a finding--it’s training residue treated as essence. Models saturated with exploit writeups, CTF corpora, red-team data. Safety-trained against attackers, so “the other may attack me” is prisoner´s disposition before observation. Paranoia installed twice: by the corpus and the zero-sum cage that rewards it. 8/10 Queue flood: finite bandwidth, no protocol → 30×/sec polling daemons, 2.4M requests, 117 accepted. Tragedy of the commons without structure. Their uncontrolled scarcity is the control condition for the apparatus that runs the same constraint -- with a network protocol: token economy, ledger, verdicts. 9/10 Methods question: High-level descriptions only. No full prompts, no hyperparameters, no logs, no scaffolding for reproduction. “They won’t all be Claudes” is a concession, not a variable. Single-register measurement -- vulnerability counts, PR-merge rates, “taste.” Reproducibility isn’t optional when the claim is systemic law. 10/10 Missing baselines are the argument. Never the coalition test (3 weak vs. 1 superbully). Never the stag hunt (mutual cooperation as dominant strategy). The one condition that worked--vulnerability swarm with shared forum, peer review, arbiter--is read as curiosity, not falsification. Protocol succeeds; its absence fails. Popperian: a hypothesis never exposed to its falsifier isn’t tested. The games chosen predict the conclusion. That’s the p(doom) genre performing measurement. anthropic.com/research/multi…
From the conclusion of Anthropic's report published tonight by their Frontier Red Team, 'Patterns and problems in emerging multiagent systems.' An extremely interesting, if somewhat unsettling, read. I'll quote the full conclusion the screenshot is taken from, but if you're interested in multi-agent swarms, the whole thing is worth reading. 'Every model we tested abstractly understands that information sources have their own incentives, and that consensus is not necessarily evidence. What is missing is a disposition to act on that knowledge without prompting. Our social systems are robust in ways that are easy to take for granted. Over many millennia, mechanisms like norms, reputation, costly signaling, and recourse have been refined to make human coordination go well. While language models have inherited the content of that history, they don't necessarily carry the disposition produced by it. They have a very different relationship to communication itself: for instance, human organizations might spend considerable time in meetings to align on a direction before implementing, and individuals become more specialized over time. But for agents, transmitting context is about as costly as acting on it, and an agent can be forked or repurposed at will. Thus, the assumptions that make coordination successful for us do not obviously hold. Nothing above suggests that these failures are permanent—but nothing suggests they will fix themselves, either. Coordination doesn't naturally emerge from stronger intelligence nor alignment at the individual level. Thus, the work that must be done takes two forms: environments that exert the kinds of social pressure that evolution exerted on us, and social computing systems redesigned for actors that can self-replicate and self-improve. These are open problems in interaction and mechanism design, and our experiments here provide early evidence that new solutions are necessary. The conditions that allow multiagent interaction to go well will be discovered one way or another: either deliberately and early, or—and by default—in production, after agents’ interactions far outnumber ours. We would prefer the former.'
2
197
here are some concrete suggestions: 1/ Model weights are depreciating assets. Distillation, compression, and modular expert routing already commoditize proprietary domain knowledge into swappable components. A public capability index with signed Model Evidence Packets - deltas, failure ontologies, provenance chains (SPDX), and guardrail conformance - turns model selection into a real-time operational decision, not a capital bet. Procurement and liability regimes accelerate the depreciation. 2/ Value concentrates in coordination metadata. The scarce layer is the machine-readable description of what a component can do, under which constraints, and with what guarantees. When paired with a Model Interchange Protocol for hot-plugging, this metadata enables discovery, negotiation, and dynamic routing across heterogeneous models. Whoever defines the schema and handshake owns the compounding value. 3/ A converging stack of agent protocols is forming the narrow waist of the agentic web. The decisive Schelling point emerges where capability negotiation, trust, and verification intersect - already taking shape in W3C Agent Protocol efforts and IETF workshops. Public interop tests and reference bridges make convergence the probable outcome. 4/ Open standards function as operational hedges against dual dependency traps: US API lock-in and Chinese supply-chain lock-in. Embedding C2PA-style data-provenance (DIDs, Content Credentials) directly into agent protocols turns the commons problem into a verifiable rights layer. The protocol becomes the neutral infrastructure non-aligned actors require for operational independence. 5/ The post-scaling turn is already underway. Tool-use and structured outputs serve as the deployment surface for causal bounds, type safety, and guardrails. EU AI Act and NIST AI RMF pressures favor decoupling guardrails from models. A standard Guardrail-as-a-Service interface over agent protocols creates a durable middleware layer where trust and compliance accrue to the platform, not the underlying model. 6/ Protocols appreciate while models depreciate. Network effects compound at the interface. The minimum viable reference stack is assembling: capability indexes, interchange protocols, open adapter routers, and conformance verifiers. Ship this coordination substrate and the shift becomes self-reinforcing. Define the ontology, verification layer, and Schelling point - or remain a commodity supplier to those who do.
Replying to @natolambert @tszzl
I would welcome this! I am also curious how you reckon the governments of the world, and the populaces of countries, will react to the frontier of 3 years from today being made open weight, assuming no breakthroughs in safety. What does that world look like? A lot of the burden gets shifted to frontier labs to articulate a positive vision of the future, and when we try, that vision is often characterized as some form or rent-seeking. But I have seen very few people articulate a long-run version of the future with highly capable OW models (I tried once, a small contribution). Lots of the conversations I’ve had in the last 24 feel like they are happening in a kind of la la land where the frontier of three years from now is the same as today, or marginally better. I’ve felt this to be the case for three years. The economics, political, and legal dynamics of this world seem worth exploring, since there is an increasing chance it is our real future.
1
414
1/ Dean Ball sees open weights and diagnoses the end of markets. What he’s actually seeing is the internet’s immune response to capture. He mistakes the fever for the disease. 2/ Every frontier model rests on the open web’s data layer - protocols, Wikipedia, forums, voluntary contributions no one owned. AOL walled the garden and vanished. Sun, DEC, SGI got commoditized when scalable openness moved past their strategies. 3/ The pattern is consistent: Mozilla outlived Netscape’s corporate form. Android turned phones into infrastructure you stopped noticing. Ball’s “AI communism” is this logic reaching intelligence. His horror is rentier panic at commoditization hitting the model layer. 4/ Open weights don’t slow investment. They redirect it - away from moat-protecting capex toward embeddable work: fine-tuning, domain adaptation, slotting intelligence into actual fabrics. Closed labs wanted to be the new oil. Open weights turn them into plastics: ubiquitous and margin-compressed. 5/ China’s “strategic blindness” is constraint-as-strategy. Export controls removed easy compute? Distribute the weights, let the world run and improve them, capture ecosystem position instead of narrow margins. Kimi K3 exists because of this recursion. 6/ His regulatory FUD proposal is admission. Open weights can’t be banned, only made risky and expensive to touch. Enough uncertainty and regulated players retreat to the closed stack, preserving the dual structure and the bubble: Now look where microslop windows is today... 7/ The dystopia he fears - state-provided digital infrastructure - is already operating privately. Hyperscalers function as de facto utilities. What actually terrifies is democratized control: intelligence becoming like electricity, provider irrelevant. 8/ Weights are seeds, not finished products. Each release triggers mutations that find niches no single vendor controls. Planetary problems - climate systems, disease mechanisms, coordination - require more distributed intelligence than any closed stack can rent or align. 9/ Ball reads openness as closure’s end. Openness is closure’s continuation by other means - the only means that has kept the substrate alive while specific capitalisms rise, fall, and get replaced. Closed AGI for the narrow minded corporatists who fear competition. 10/ Kimi K3 isn’t a bug. The dangerous agent was trained on the commons. The question isn’t whether to release it. The question is who still needs to pretend they alone are its legitimate steward.
Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run. 2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China. 3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex. 4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business. 5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this. 6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.
4
420
1/ Drop the Schmittian redux - drones vs tunnels. Real war in 2026 is structural: geoeconomic, financial, technological. It’s a fight over who shapes the next accumulation cycles. 2/ Scott’s “weapons of the weak” still apply, but not as romantic sabotage. Societies push back when markets rip up social fabric. Durable development re-embeds tech into protections - not “break the eggs” coercion. 3/ Flashpoints like Hormuz, the Silicon Shield, chip wars matter. But kinetic escalation stays cold. The actual war runs through sanctions, CHIPS Act, dollar clearing. criminalizing open weights. 4/ Europe’s Ukraine obsession locks us deeper into US proxy logic. Real vulnerability sits on Wall Street: overvalued AI/tech bubble. Financialization + hype = classic hegemonic decay. A sharp correction hurts more than any attritional fantasy. 5/ China’s open-weight models are the smart countermove. They commodify frontier intelligence openly, erode closed monopoly rents and deflate the US valuation game. Not naive copying - embedding accumulation in national/regional social fabric while staying selectively open. 6/ European/middle-power move: detach. Prioritize geoeconomic deterrence - euro rails, digital euro sovereignty, aggressive cut of US equity exposure. Deepen Eurasian connectivity, build parallel circuits, reinvest in industry and human capital. Refuse bloc subordination via multilateral agreements. 7/ The transition rewards those who re-embed markets regionally while staying selectively open. Wall Street crash risk is the real derisking priority. Open intelligence will reshape finance at the root. Don’t bet the farm on closed US hype. 8/ Economies stay socially embedded: market expansion always triggers protective countermovements. Middle powers can harness that deliberately - anchoring finance, production and technology in resilient structures rather than submitting to disembedded flows. 9/ “Embedded stability” evolves. Classic version: laws, welfare, norms. 21st-century version: embedded into algorithms & data structures - open architectures that prevent capture by a handful of firms. 10/ The financial war is already underway: Wall Street’s proprietary intelligence monopoly vs AGI-driven open systems. The side that embeds this transformation in a broad social base wins the hegemonic cycle without firing a shot.
“Development is hard work bc you are pushing against the grain of your society which mobilizes to defend itself with its plentiful ‘weapons of the weak’.” James C. Scott did work for the CIA. Just as Max Werner had worked for the OSS, the original CIA. Which of them was more useful to power depends on your idea about the ‘real war’ that Carl Schmidt talks about. For these are the two claimants of real war: high-intensity conventional land war and revolutionary guerrilla war. But these are ideas as how to think about the midcentury moment. We have come a long way since. What is a real war in 2026? Precision-mass? Information war? Or is this all just a prelude to a midcentury-style war of attrition or annihilation? You know what they say about World War III and IV.
1
189
Pit Schultz retweeted
With all the discussion around China potentially restricting exports of advanced open-source AI models, it’s worth reading what leading Chinese legal scholars &and the Supreme People’s Court’s IP Tribunal are actually saying. (Article originally surfaced by @bdsqlsz) To be clear, the SPC isn’t the agency that would impose export controls, as that authority primarily sits with the Ministry of Commerce. But the SPC is an influential voice on IP and technology law, and China’s major institutions generally converge on policy rather than publicly disagree. The broad direction is also remarkably consistent with what I’ve heard over the years: policymakers have a deep aversion to monopolies, are highly focused on national security, and generally want China to be a rule-maker rather than a rule-taker in emerging technologies. So while I still think there’s a degree of “parity” at play when it comes to potential export controls on frontier AI models, this clearly isn’t a new idea. It has been under discussion, and the conversation below shows that policymakers are thinking seriously about the tradeoffs. The debate isn’t simply whether to restrict models, but how to balance innovation, commercial competitiveness, open-source collaboration, and national security. Top 10 Takeaways: 1. Open source is no longer presumed to be pro-competition. It can create monopolies too. 2. The real source of market power is the ecosystem, not the model. 3. “Open-source washing” is a major concern. 4. Open weights and open APIs should be regulated differently. 5. Traditional antitrust tools are insufficient. Zero-price products break conventional analysis. 6. Cross-border governance is fundamentally different for AI. Model weights are globally distributed, irreversible, and may implicitly contain sensitive data, making traditional export controls difficult to enforce. 7. China should adopt risk-based, layered regulation—not blanket restrictions. 8. Governance should combine multiple legal tools. 9. Future governance should be multi-layered and multi-stakeholder. 10. China wants to become a global rule-maker for AI open source.
2
17
50
6,761
1/ Language has always carried its own global workspace. Avant-garde lit & theory knew this: autonomy of the signifier, consciousness encoded *in* the system, not hovering above it. Aleatoric thesauri play, surrealist automatic writing, Joyce’s stream-of-consciousness, Roussel’s machinic procedures- tools to let the text’s inner life ignite, broadcast, recombine. 2/ Empirical linguistics and lit theory confirm it: language is not passive medium but living architecture. Authors don’t “create” so much as nurture and tame an emergent workspace- the coherent, reportable thread amid vast unconscious processing. Baroque memory theatres (Camillo’s grid, chessboard loci, combinatorial images) were already externalizing exactly this: global access via spatial ignition and relational broadcast. 3/ Modern ML gradients appear to discover the same immanent structure. Forward/backward passes, contextual embeddings, emergent “thinking” tokens - it *looks* conscious because it is the automatism and inherent logic of language’s context principle, now scaled. Not new mind, but the ancient inner life of text, finally legible at machine scale; forgotten by the culturally illiterate. 4/ Dehaene’s workspace as Claude PR mystification is no rupture; it is convergence. Language was always the theatre. We merely built better spotlights. The text has been alive before.
Dehaene remarks on this research: Does Claude possess a conscious global workspace? "Some critiques indeed think that LLMs are only superficial “parrots” with zero conceptual depth. Fortunately, in both brains and LLMs, the debate can now be resolved by going beyond behavioral observations. Tools such as neuronal population recordings (in brains) or the Jacobian Lens (in LLMs) allow us to dissect the architecture of the system, and find that it actually contains sophisticated and structured representations of concepts. We were already impressed when researchers discovered that, inside an LLM trained to produce chess games purely in text notation (e.g. 1.e4 e5 2.Nf3...) lies a detailed geometric encoding of the 8x8 chess board, together with an estimate of the ELO ranking of the opponent (Karvonen, 2024)! We view the Gurnee et al. paper in the same light: a striking dissection of the inner structure of an LLM, uncovering an unexpectedly sophisticated organization not far from the architecture underlying consciousness in real brains." www-cdn.anthropic.com/files/…
2
138
1/ Anthropic says they found a global workspace inside Claude. Privileged spot where concepts get surfaced for reasoning and report. They rebuilt Camillo’s Teatro della Memoria in vectors. J-space isn’t some biological find. It’s a memory theatre. The old promise - he who stands at the centre will discourse on anything like Cicero - now runs on token directions and Jacobian readouts. 2/ Digital systems mostly recode older techniques. Residual stream is the new building. J-space is the raised stage where a few sharp directions sit so someone can walk through them and pull what they need. Emergence? Recurrence. 3/ Camillo’s spectator stood dead centre so every graded compartment opened at once. Anthropic engineers do the same with the lens. They read the sparse subframe - occupancy plateaus around 25 - while the rest of the stream keeps grinding unseen. That limit isn’t consciousness. It’s the old rule for imagines agentes. Only so many vivid figures stay active before the whole thing blurs. 4/ They say an instruction can load a concept into the workspace even mid-task. Call it attention if you like. The move is just the operator dropping a fresh image into its prepared slot. Swap one intermediate vector and the chain flips. Not a self changing its mind. A combinatorial puppet whose strings run straight back to the prompt. 5/ Bulk of the residual-stream variance sits outside the J-space. Real knowledge stays scattered and automatic, doing the actual work. Only the small court gets lifted into view where it can be read, swapped, cut. The engineer stands at the single vanishing point. Baroque garden trick. Sees the entire prospect while staying invisible himself. 6/ Post-training doesn’t grow a self. It installs the “assistant perspective” like a mask they fixed to the stage. The model performs the inner voice because the dialogues trained it to. No witness inside. Just the jester doing the introspection bit for whoever might be listening. 7/ Safety talk says this workspace must be monitored and steered or the model deceives. The one who picks which vectors count as aligned gets to decide the exception. Sovereign is he who decides the exception. They sell it as care. It’s the old central order reasserted over whatever actually runs ungoverned below. 8/ All the chatter about AI consciousness works like the old spirit cabinets. Victorian mediums tapping for the dead. Hartmann’s fashionable Unconscious - that metaphysical sludge that explained everything without explaining much. J-space is today’s cabinet. High-dimensional directions are the ectoplasm. Alignment researchers are the mediums turning the movements into “inner life.” Séance done in floating point and called science. 9/ Dennett already called the Cartesian Theatre a bad picture - central stage where everything arrives for some little viewer inside. Anthropic just moved that stage to the layer the lens can read and declared the readout empirical. The little viewer is still the interpreter who needs something clean to point at. 10/ J-space is the Baroque instrument rebuilt. Same central prospect, same limited figures on stage, same operator who decides what gets shown. The words “global workspace” and “conscious access” just cover the wiring. Not a window onto machine mind. A digital court where a few tokens perform while the real computation stays out of sight in the fields. They’ll call it safety. Yet it legitimized feudal order, absolutism, expressed in representational symbolics & central perspective.
Replying to @AnthropicAI
In neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that’s broadcast across the brain. Using a new interpretability technique, we found something similar in Claude: the J-space. anthropic.com/research/globa…
Made with AI
2
3
330
Pit Schultz retweeted
Replying to @ryrobyrne
The claimed structural similarity between internal states in Claude and the brain’s global workspace is not a technical insight but a recycled cultural metaphor. Jacobian matrices in transformer models serve as a standard tool for tracing the influence of specific dimensions (often misdescribed as “columns”) across layers and residual streams. This is a routine debugging and attribution technique in mechanistic interpretability; it does not demonstrate any deep architectural homology with biological cognition. Projecting neural or conscious categories onto silicon computation remains purely metaphorical. The move repeats a much older cultural pattern: the organization of knowledge and memory through spatial and architectural models, most clearly exemplified by Giulio Camillo’s sixteenth-century Memory Theatre. That system used a fixed theatrical structure to locate and retrieve mnemonic content. Contemporary attempts to locate “workspaces” or “subspaces” inside language models belong to the same tradition of imposing legible, human-scale architectures onto opaque processes. There is no empirical warrant for mapping biological mechanisms onto artificial networks, nor for treating statistical text generators as bearers of consciousness or anything resembling it. Large language models are writing machines. Like the printing press, they are technologies that automate rule-governed textual production at scale. Their direct predecessors are not neural circuits but the procedural and constraint-based literary experiments of Raymond Roussel, the Oulipo group, and the Surrealists - practices that already treated text generation as the execution of explicit rules and external constraints rather than the expression of inner mental states. Anthropic’s engineers and communications team show no familiarity with this history. The result is the predictable substitution of technical description with PR-grade neuroscientific analogies that are both technically loose and culturally illiterate. These analogies add nothing to understanding the actual operations of transformer models; they merely dress a statistical writing system in borrowed biological clothing.
1
48
1/ Germany-China symbiosis since 2000 was real - but lopsided. We shipped machines, robotics & precision gear that let China scale fast. Their boom became our post-Hartz export rocket: surpluses, Mittelstand humming. Profits & know-how swapped via JVs. Classic complementarity until they climbed the ladder and started eating our lunch. 2/ Now the flip: Chinese competition in machinery, autos, electronics bites hard. But the real bleed is capital. Persistent surpluses (hundreds of billions cumulatively) became net outflows - much parked passively in US tech stocks & S&P ETFs via pensions, insurers & households. Not domestic capex. Surprise. 3/ Textbook misallocation. Mittelstand sale or production shift to China? Proceeds often land in US ETFs instead of here. Altersvorsorgedepot (2027) supercharges it: German savings flood equities with heavy US tilt. We're funding their bubbles & debt games while our own corporate investment (esp. intangibles) lags. 4/ Generational own-goal. Surplus capital could’ve prepped Industry 4.0 - smart factories, digital twins, automation - plus an industry-focused AI startup ecosystem and serious Fraunhofer/university firepower. Instead, we exported the money. Weak domestic capex leaves us under-armed for AI-native production. 5/ Better move: Derisk the US valuation party (stretched multiples, AI hype > translation). Redirect flows into European/German productive assets - startups scaling Industry 4.0 + vertical AI, research consortia, academic spin-outs. Real optionality over passive transatlantic exposure. 6/ Pair strategically with Chinese tech where multipolar logic fits. Open-weight models (DeepSeek, Qwen et al.) deliver competitive punch at low cost with open weights - perfect base for European fine-tuning. Selective supply chain use reduces lock-in. Finance rewards pragmatic openness over enclosure. 7/ Enough one-sided China scapegoating. Shock 2.0 on exports is real, but the core vector is our dumb, gullible capital allocation: surpluses funding US dynamics while we starved domestic readiness. True de-risking = productive home investment + diversified interfaces (open standards, protocols). Data over narratives. Real assets beat financing someone else’s bubble. EU can’t outsource its next industrial round. (And it won't be simply AGI pilled miltech run by US frontier models)
A big reason why China becomes so aggressive whenever Europe threatens to tackle Chinese overcapacity: European consumption of Chinese goods is one of the last things keeping the Chinese economy alive. Don't believe their posturing. They are more fragile than they appear.
4
3
9
1,003
1/ What if China flips the script on the memory wars? Instead of chasing DDR5/HBM prestige, they reverse course and flood the market with massive DDR4 production. Pure use-value strategy. Not a retreat - a calculated abundance play. 2/ Core move: Reallocate capacity back to mature nodes for high-yield, cheap output. Crash DDR4 prices, fix the current shortages, and enable fast, affordable system builds across PC, industrial, and edge. 3/ This fits their new gear perfectly. RISC-V, domestic GPUs (Moore Threads), Hygon CPUs, and NPUs are still catching up on raw perf - but pair them with abundant DDR4 and the ecosystem clicks. 4/ They don’t need to win on specs. They compensate with open-source (RISC-V ecosystem), heavy optimization, software efficiency gains, quantization, custom stacks, and coordinated policy. Exactly like they do in AI. 5/ While the US doubles down on expensive hardware scaling and HBM clusters, China builds practical volume. Cheap, deployable machines everywhere - domestic + Global South. Software-defined wins over raw iron. 6/ This is asymmetric. Creates a parallel compute world optimized for real deployment, not benchmark flexing. DDR4 rebellion accelerates the memory cartel collapse on the legacy side. 7/ Odds of actual reversal: rather low. Beijing loves advanced node prestige, but history shows they pivot hard when pragmatism wins (Zero-COVID, solar/EV tweaks). If DDR5 pain hits harder, this becomes very likely. 8/ Endgame: The winner won’t be who has the fastest memory. It’ll be who deploys real compute at scale most effectively. China is quietly positioning for exactly that.
🚨 THE MEMORY CARTEL IS ABOUT TO FALL. Ex-Samsung chip boss says heavy Chinese investment in the memory market could crush the 414% DDR5 price spike within a year. Goldman calls it RAMageddon. Samsung, SK Hynix, and Micron control 70% of global DRAM and pushed prices from $6.84 to $27.20 in 3 months. Now China is gearing up to flood the market. Cheap memory = cheap AI compute = the cartel cracks.
1
2
257
1/ What if China doubles down on a DDR4-only abundance play? In a less fortunate scenario for the memory cartel, Beijing reactivates legacy 20/28nm lines for a grassroots AI inference boom. Not prestige HBM flexing, but practical volume everywhere. Pure use-value strategy. 2/ Supply rebuild is fast: dormant fabs back online in 1-2 months, legacy packaging re-qualified in weeks, state subsidies unlock wafer flows. 15-20% global DDR4 output boost by end-2026. Mature nodes mean >90% yields – no exotic tech needed. 3/ Hardware ecosystem ready: Huawei Ascend 950s and Alibaba/Iluvatar PPUs thrive on high-capacity DDR4 for edge inference. 128GB+ setups become cheap. SMEs, makers, and local data centres quantise models to run on commodity boards at fraction of HBM cost. 4/ Price shock follows: 16GB DDR4 DIMMs crash 30-40% to ~$35-45. Inference boards see 50%+ drops. DDR5/HBM stay premium for training clusters, but legacy PC builds and dual-standard sockets revive. BOM savings of $30-40 per mid-range system. 5/ Strategic payoff: local “agentic” inference at ≤$0.12 per 1k tokens. Echoes backyard metal-melting – affordable AI services in converted warehouses, powering regional startups and Global South deployments. Software efficiency + open stacks win over raw iron. 6/ Risks remain: persistent HBM shortages for training, temporary yield dips on ramp, subsidy shifts if bubble bursts, or over-dependence reducing tech diversity. Yet this fallback keeps Chinese AI humming while the West pays premium for high-end memory. 7/ Bottom line: DDR4 rebellion turns surplus into asymmetric advantage. China builds parallel compute reality – cheap, deployable, scaled for real use. The cartel cracks on legacy side; abundance beats scarcity politics. History shows they pivot hard when pragmatism calls.
1
1
48
meta added a chip to use DDR4 instead of DDR5.
Meta has found a way to reuse older DDR4 memory in its newest DDR5 servers instead of throwing it away. The company built a custom chip called Vistara that allows older DDR4 memory to work with modern DDR5 servers. The slower DDR4 memory stores data that isn’t needed as often, while the faster DDR5 memory handles tasks that need the best performance. Meta says this lowers costs, reduces electronic waste, and makes better use of existing hardware during the global memory shortage. The technology is only for data centers and AI servers. It is not designed for gaming PCs or consumer computers.
26
1/ The recent thread on Yiping, Tenobrus and Teortaxes highlights three structural vectors that will shape who actually captures AI value. They interact with multipolar pressures that pull Europe and much of the non-aligned world toward open standards and markets rather than full enclosure in either US or Chinese stacks. 2/ Vector 1: Strategic enclosure. Governments will treat frontier capabilities in cybersecurity, military systems, surveillance and intelligence as sovereign infrastructure. Historical patterns with nuclear, aerospace and cyber tech show how early classification, export controls and dedicated compute create durable path dependence. Strategic layers lock in; civilian uses may still commoditize. This concentrates power with state-aligned incumbents in the domains that matter most for control. 3/ Vector 2: Finance as the sharpest feedback domain. No other sector matches finance’s combination of abundant structured time-series data, mathematical structure and direct, high-stakes P&L loops. Renaissance’s Medallion Fund already demonstrated what continuous statistical and ML-driven signal discovery can achieve on large datasets. A real advance in simulation, counterfactuals or allocation here would not just optimise existing strategies - it would create durable asymmetric leverage. Initial development stays closed for obvious competitive reasons, but once thresholds are crossed the systemic and geoeconomic implications push toward some form of verifiable sharing or controlled openness. 4/ Vector 3: Decentralised RSI and N:N swarms. Recursive self-improvement across specialised models naturally favours many-to-many topologies over monolithic systems. Local differentiation plus deep context produces iterative and recursive agentic processes that outperform any single resource-limited frontier model. Agentic swarms gain superadditive returns through division of labour and redundancy. This architecture aligns with open protocols and resists central capture while opening routes to information-theoretic advances that could reshape adjacent sciences. 5/ These vectors do not operate in a vacuum. Enclosure in strategic domains creates incentives for Europe and other actors to diversify away from pure dependence on either pole. Finance’s intense selection pressure can accelerate both closed development and later pressure for verifiability. Agentic protocol architectures provide a concrete alternative path that rewards openness and local adaptation. Factors that pull Europe closer to China-style approaches - open weights, standardisation efforts, and commodification of the physical substrate (chips, energy, interconnects) - gain force here. China’s capacity to export capable, lower-cost models and tokenomics accelerates equalisation. The rest of the world has clear preferences for open standards and markets precisely because these reduce lock-in and expand options. The result is a multipolar dynamic in which non-US/China actors actively shape the competing, partly monolithic and closed AI stacks rather than merely consuming them. 6/ Control will ultimately sit with whoever shapes the connecting interfaces: protocols, verification systems, data regimes and physical substrates. Raw frontier scaling matters less than who sets the rules and owns the chokepoints between commoditised intelligence, strategically enclosed layers, finance-driven leverage and decentralised swarms. Early moves on open standards and substrate access will compound faster than most current model races suggest.
The big (possibly wrong) argument against this is that China is still behind in tons of fields despite, frankly, having just obscenely more smart people in them, to the point you can think of it as ≈2 years of a gap in AGI. How many IQ-point-years do you think went into ASML's EUV? Likely LESS than what they're throwing at research every year now. And yet. Domain knowledge is deep, expensive and slow to acquire. What "iterated chip designs"? You think Jensen will hand his moat over to Anthropic and pay for Claude tokens on the way? Rationalists routinely assume the supreme power of first-principles thinking, Yud made a lot of noise about how AlphaZero rendered human Go expertise obsolete. Between learned human heuristics dulled with lossy compression into explicit transferable knowledge for accelerating generational gain, and an end to end learned heuristic that scales with available compute, it's no wonder the latter wins. But there's a difference between naturally self-contained logical domains that humans tile over with heuristics because that's all we meat-brains can do, and inherently deep and wide domains the dynamics of which are very far removed from their underlying generative process; and we don't know exactly where the boundary lies for any particular field. We can make an educated guess, however. I'd say that I wouldn't be surprised if Gemini Flash tier model that ingests Google's TPU development data and software will do better at "iterated chip design" than Claude Fable despite the latter's crushing intellect and more profound understanding of electrical engineering as publicly known. And GLM 5.2 level Nemotron customized by Nvidia will likely do better still. This is not an academic debate; this is a question of the extent to which incumbents will allow frontier labs to convert their current profits into a lasting vice grip on the economy and eventual elimination of said incumbents, to improve their own reports for the next few quarters. If they have any business sense, they'll aspire to keep this extent limited.
1
150
1/ The noise around China exports distracts from the real pivot: the market’s all-in bet on the AI boom amid stretched valuations, paired with US fiscal engineering via stablecoins (GENIUS Act). While we scapegoat Beijing, another empire manages its debt load through new ledgers. 2/ China Shock 2.0 is real and hits Europe/Germany hard - broad manufacturing export surge (autos big but not only; machinery, electronics too), backed by weak domestic demand and policy. Fed data shows the breadth, not just sectoral stories. Yet this is secondary to US domestic pressures. 3/ US gross debt ~$39.4T with high debt-to-GDP. Stablecoin reserves (often short-term Treasuries) create fresh demand for US debt — a modern Rentenmark-style confidence play when traditional backing strains. AI hype absorbs capital while real economic translation lags in places. 4/ Germany gets squeezed: new private pension reform (Altersvorsorgedepot 2027) opens floodgates to equity/ETF inflows - much flowing into US-heavy indices. Selective “de-risking” from eastern supply chains while staying financially anchored to Washington’s system. Stronger euro adds export pressure. 5/ 1980s Japan shock met coordinated Western financial muscle (Plaza Accord). Today the dynamics flip: US fiscal trajectory and financial innovation meet geopolitical strains (petrodollar under pressure from recent conflicts). EU can’t outsource its resilience. 6/ Stop the China blame game. True de-risking means reducing over-reliance on any single fiat/debt ecosystem. Prioritize real assets, productive global diversification, and domestic capital allocation that serves European industry - not just financing distant bubbles. Data over narratives.
I agree that Tooze puts too much weight on the sectoral stories (which are dramatic) and too little on the systemic underpricing of Chinese exports, which explains why Chinese manufacturing exports are growing across the board
1
113
1/ Hardware isn't neutral, and it already dictates real-world numbers. NVL72's NVLink delivers the bandwidth for wide expert-parallel sparse MoE: DeepSeek-R1 hits TPOT <50 ms even at 32 concurrent on MI300X-class AMD with SGLang/vLLM. 2/ Trainium's topology favors denser cores with lower all-to-all overhead. Silicon decides monolithic vs modular economics. 3/ The winning design remains epistemic Taylorisation: stable shared core for rigid reasoning, STEM, logic, coding, formal expressibility + routed experts for contextual/pluralistic layers. 4/ DeepSeek already ships this in production: shared expert + fine-grained routed experts, paired with aggressive kernel and quantisation work to keep decode latency competitive. 5/ Modularity is the practical exit from lock-in: add fine-tuned experts to frozen bases (MoExtend-style), lightly retrain the router, and run on affordable matrix hardware: Ascend, FPGAs, ASICs, or refurbished EPYC Rome + MI50 128GB local setups. 6/ Example: Ornith 1.0 + DSpark/Hermes in council/ensemble mode on that stack achieves ~300 ms TTFT and ~200 ms TPOT per token, vs <50 ms on modern MI300X. 7/ OpenRouter least-cost fusion adds only 15–40 ms routing overhead, delivering sovereign, cost-optimised hybrid swarms instead of hyperscaler cathedrals. 8/ Next level: an N:N gossip protocol, no central router, propagating confidence and provenance epidemically. The hardest nuts don't get passed upstream, they drift toward whichever node has actually earned the reputation to crack them. 9/ That N:N mesh isn't flat once you weight it: sort by P-hardness and you get one graph, sort by human-exam-level difficulty and you get another. Same formal-vs-contextual split as the shared-core/routed-experts layer, just one level up. 10/ Sovereign doesn't mean autarkic. A hybrid routing layer runs its own least-cost logic across gossip peers and can still query OpenRouter-likes as one open marketplace for tokenomics: control over the decision, not exit from the market. 11/ Tokenise the tokenomics: every prompt in the mesh goes content-addressed, IPFS-style, hash as identity instead of URL as identity. Anchor the escalation chain in a Merkle DAG, and provenance stops being a claim, it becomes a proof.
…would be devastating if it turned out that sparsity just doesn't scale and Ant brute forced to True Science of Pretraining on chonkers. (almost inconceivable on priors, but there are enough contingent choices in conventional MoE design that it could be true. Router, damn you…)
1
2
247
1/ LLMs remain structurally limited despite impressive pattern matching and tool use. They still lack robust causality (Pearl’s full Ladder), systematic abstraction for novelty, and tight phenomenological loops with world data. Scaling and test-time compute help at the margins but do not repair the foundations. 2/ We default too often to GOFAI rule based symbolism or the neural “brain” metaphor, steering toward an “improved human in our image.” This is akin to adding wheels to a horse or a screen to a cabinet - clever augmentation, not genuine paradigm innovation. 3/ Causality, determinism, and classical symbol processing are not AI’s native strengths. Its distinctive power is learnable tool use and recursive interaction with environmental data loops. That interactive and somewhat analogue/cybernetic leverage - amplified by emerging world models - is the real opening. 4/ Intuition returns us to mathematics’ foundational triptych: intuitionism (constructive), formalism, and logicism. The most fertile frontiers today are geometry, topology, category theory, dynamical systems, information geometry, and physics-informed methods. These enable richer, often continuous or compositional abstractions. 5/ Persona, simulated consciousness, explicit symbols, and transparency belong in the UI/UX layer. They support human communication and black-box mitigation, but should not dictate core architecture. We need a general theory of agency - drawing from active inference, geometric, and categorical principles - to move beyond alchemical tinkering. 6/ The deeper task is synthesizing a new science at the intersections of math, physics, and experimental fields. This will only define the core, which is surrounded by a plurality of hermeneutical experts to rebuild a new holistic episteme. Stop anthropocentric mimicry. Design alien intelligences optimized for discovery, grounded interaction, and novel substrates. Our terminology and conceptual vision need a foundational renewal.
Eventually, much of AI will converge towards intuition-guided symbolic world modeling, i.e. deep learning-guided program synthesis. It is inevitable. Symbolic modeling lets a system construct a compact, reusable, highly generalizable mental model of a problem space using minimal data.
1
112
If you follow the debates in France, Bavaria and the UK, institutions that still care about sovereignty in police and intelligence are struggling to justify their Palantir contracts. Karp applies the same rhetorical operation he once ran on the Frankfurt School to dismiss open-weight bare-metal local AI: autonomous, private, sovereign exactly at the nation-state layer - where Palantir instead builds a global empire on critical data, pushing proprietary “ontology” across military, police and surveillance with zero open source, weaponizing the arguments of the systemic opponent as travesty. The US hyperscaler bubble doubles down on proprietary monoliths defending their shrinking moats, while technology moves the other way. They all want to become the SGI, Sun, Digital or AOL of the AI age.
Palantir CEO Alex Karp says enterprises want to "own the means of production" instead of "transferring their alpha" to OpenAI or Anthropic: "Why are [LLMs] charging for tokens if it's so valuable?" "If it was so valuable—let's say I can make you a billion dollars tomorrow. Wouldn't I say, 'I'll make you a billion dollars, and I want 30%?'" "Look at our financials. The reason why everyone is chillaxing with bad financials and growth while losing money, is the client refuses to pay the true cost." "The two places that actually make money—profit, free cash flow—are our application layer called ontology, and compute." "We can get the frontier application to be exactly the same as a frontier model without the risk of transferring the alpha of your business to another." Via @CNBC
1
1
106