"Turns out Awe in its original sense meant extreme reverence and dread, derived from the old norse word “agi”"
Was looking up the best term to describe the sublime ambivalent mix of excitement and terror “holy shit it just did in minutes what would have taken weeks across a team” Turns out Awe in its original sense meant extreme reverence and dread, derived from the old norse word “agi”
1
88
We need a SKILLS_PATH so that we could distribute skills like any other libraries and not copy them into every project using them.
2
100
When you install software it just needs to be 'in the path' to be usable in a project (in the executable path, in the library path - etc). But skills need to be copied into the project that uses them. This makes it difficult to package skills that would be part of a bigger system and reusable across harnesses. I know that I can package everything into the skill - but then I get duplication if I have a few skills that use the same components of if I have a few projects that use the same skills. In general software distribution has been there for so much time that it covers really many diverse usage cases - I think it would be easier to make skills join that model than adjusting the skills model for each case.
49
Think about two way an coding agents say could propose adding a new field to an API: - expose the value - surface the value Which one would be accepted more often by human evaluators? This is why llm written text hides warnings and is so hard to evaluate. RLHF is the root of these problems.
34
zby retweeted
Please repost the crap out of this until @elonmusk or @Gwynne_Shotwell replies to it: Grokipedia is dead. I'd like to rescue it, starting with scientific data, which is critical to intelligence (artificial or not) ... 1. LLMs have a GIGO problem. They ingest tons of wrong information and spew it out with certainty. 2. Peer-reviewed research is a disaster. It's a $10 billion industry with bad incentives and terrible results. Feynman would not be happy. 3. Far too many global crises are caused by fake, weaponized, political, and simply incompetent science. I believe the ONLY way to fix it is not with programming and training, but with an open, two-sided market mechanism where you make money if you're right and you lose money if you're wrong. It's not a prediction market, it's a validation market. This skin-in-the-game approach will produce the world's most accurate science data, and much more. It's better than community notes. It could become the basis for an entirely new media company that complements X.com. DM me and I'll send you my deck. Sincerely, David Siegel @chamath @bgurley @hthieblot @jefielding @larjo280 @naval @pmarca @Jason @paulg @bhorowitz @balajis @fredwilson @msuster @Jeff @hunchventures @jasonmendelson @sether @jeremysliew @VCMike @ChristopherA @jeffnolan @ryanhoover @hstebbings @garrytan @alexf @jaltman @jasonlk @jessicalessin @joshua @khoslaventures @nicole @petersuciu @BillAckman @saranormous @JeffDean @kristianfreeman @alexwg @daveblundin @PeterDiamandis @kcoleman @matthew_pines @shaunmmaguire @dwarkesh_sp @C_Angermayer @RichardSSutton @deedydas @ID_AA_Carmack
18
26
184
66,214
Peer in „peer review”
1
29
The fact that the overload of the review system becomes evident on the same day as the full arxive data pack is published.
26
Do you people ask ai agents if an online business is legitimate? I wonder what will be the end results if everyone does this. Here is a report on a twitter advertisement: I would not treat Miller & Hill as an established, trustworthy retailer yet. I can’t prove it is a scam, but there are enough warning signs that I’d be cautious about ordering. The strongest signals I found are: The domain is only about one month old. millerandhill.com was registered on 14 August 2026. Despite that, the site already displays multiple “verified buyer” testimonials and the Uffington cardigan shows 29 reviews. That is possible for a new store, but given the domain age I would not assume those reviews are independently verified. Almost everything on the homepage appears to be permanently “50% off”, including the cardigan, trench coat, blazer, pullover and shirt. That’s a common marketing pattern among dropshipping stores rather than evidence of an actual previous retail price. The company claims to be a US company and gives 4008 Laramie St, Suite D, Cheyenne, Wyoming and +1 651-560-7579. But that same phone number and address are also used by another newly created clothing store, Daisy & Dream, whose privacy policy identifies the business as Runaway Brands LLC. Daisy & Dream’s domain was itself registered only on 4 August 2026. This strongly suggests a company launching multiple storefront brands rather than Miller & Hill being the long-standing fashion brand the presentation might imply. The physical address isn’t obviously a clothing operation. BBB records show another unrelated business, First Command Financial Planning, at 4008 Laramie St, Cheyenne. That doesn’t prove anything—Suite D could be a registered/virtual office—but it makes the address less reassuring. Their terms mostly identify the contractual party simply as “Miller & Hill,” without giving a clear legal-company name, and the contact section gives only the support email. Their returns policy says you pay return postage and must email them to obtain the return address. For an international order, that can become important if the actual return destination is expensive. So my assessment would be roughly “probably a real Shopify seller, quite possibly a dropshipping/private-label operation, but too new and opaque to consider low-risk.” I wouldn’t conclude that they’ll simply steal the money; the more plausible risk is receiving an inexpensive garment different in quality/fit from the polished photos and then discovering that returning it is inconvenient or costly. If you really like that cardigan, only pay with PayPal or a credit card with chargeback protection, not bank transfer/debit, and I would assume the apparent $149.90 → $74.95 discount is marketing rather than a genuine 50% bargain. I also searched for the Uffington cardigan outside Miller & Hill and couldn’t find an established independent retailer or history for that exact product name, which makes me somewhat more cautious.
1
130
meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can. everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning? we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched. how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance. we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst WIRED has the first external account of the company and the work: wired.com/story/russian-star… full writeup, the setup, and all the numbers: mostik.ai/read-more
295
259
2,443
1,273,482
I don't feel much resonance with any current political theory or large-scale political movement. They all seem deeply flawed to me, though usually in ways their internal epistemology makes unusually difficult for them to see. More fundamentally, I am suspicious of “political capabilities research”: the search for better methods of converting latent grievances, identity cleavages, and coordination failures into organized power. Think Pol Pot, quite literally. But also Lenin’s vanguard party, fascism converting national humiliation into a totalizing coalition, Maoism identifying the peasantry as an enormous but undercoordinated political base, or modern populism turning diffuse distrust of institutions into a leader-centered movement. The grievance need not be imaginary. Often it is completely real. The dangerous innovation is discovering a neglected partition of the human ecosystem, such as workers versus owners, rural versus urban, colonized versus colonizer, believers versus secular elites, renters versus homeowners, and then supplying that population with an ontology, an enemy, common knowledge of its numbers, and a mechanism for coordinated action. An ideology is simultaneously a world-model and a coalition-building technology. Quite akin to an agent swarm, the resulting political group can acquire power and interests of its own, distinct from those of both its members and its enemies. It may harm the outgroup through repression while harming the ingroup through conscription, purity spirals, censorship, economic sacrifice, and demands for loyalty. Once the coalition exists, its revealed objective often shifts from advancing the welfare of its constituents to preserving and expanding the coalition itself. This is the political capability externality. Cynicism aside, the typical pattern in political innovation does seem to involve searching for a kind of power arbitrage: finding people whose preferences are important but poorly coordinated, then organizing them into a political force more aligned with the interests and aesthetics of the framework’s creator. Sometimes this produces genuine liberation. But even then, the movement’s model of reality has been selected partly for its ability to recruit, coordinate, and defeat rivals, not merely for being true. No political movement I know of seems aware of its own selection effects, epistemic distortions, and emergent power-seeking tendencies to anything like the degree that would make me comfortable entrusting it with the future. The fragile and apparently declining international equilibrium of overlapping powers, jurisdictions, institutions, and cultures therefore seems important to preserve. It is inefficient, hypocritical, and frequently terrible. But it at least prevents any one political ontology from obtaining irreversible control over the whole future. It leaves ecological niches in which alternative values and ways of life can survive. More importantly, it preserves the possibility that Team Consciousness survives our political experiments, rather than being permanently captured by whichever coalition technology happens to win first.
14
10
77
4,341
I am presenting "Where It Lives Is Not What It Is: An Architectural Vocabulary for Retained Adaptation in Agentic Systems" at ASISAS Mon 7 Sep 2026 14:40 - 15:00 at B1.2.21 (Lecture Room 21 at the second floor) conf.researchr.org/program/e… This is the first result created from my notes at zby.github.io/commonplace/
2
67
When using LLMs to write plans or prompts for execution by other agents ask it to use Auftragstaktik. I am working on a theory why this works - in short preview: war is a good advanced analog field for agentic work, because the predictable is covered by traditional software and what is left for agents is the unpredictable residue. Plus Auftragstaktik is a very characteristic word pointing to a methodology created for dealing with similar circumstances - it precisely activates the vast knowledge the LLMs have in their weights.
2
63
I am of the opinion that many of the best engineers become the best not due to their academic education but by engaging in these hobby fields early on. Those that pump out Project for the fun of it will always be better then their peers going in the field for money. Smarter governments support these hobbys and do not try to restrict them beyond reasonability. Doing this in Germany lands you in 6 figure fine territorry @xjet is a channel you should follow btw
Chinese hobbyists are experimenting with rockets on model aircraft.
38
261
3,475
178,714
youtube.com/watch?v=xH7U7w9Q… Sutton loughs off the suggestion that systems can learn outside of weights! I am surprised - everything else is so spot on - but they talk as if the llm model was the only part that can do learning. Later they admit that the state can be in other parts of the system - but stick to the idea that true learning is in weights. Also they claim that there are no cases where the system learns a world model and then plans using it. My contention with this is that I do that all the time with Commonplace - then they would say that the ability comes from me not from the computing system - but I believe that the part that comes from me is smaller and smaller - I need to work less and less with my system. I am not there yet - but ...
1
53
The whole discussion on agent swarms and other multi-agent solutions is muddled by not distinguishing two things - one is parallelism - the other is using clean context.
48
a continuation of my neo-slop complaints: 4) subsystems responsibilities overlap 5) solve problems at the wrong architectural layer 6) duplicate sources of truth, then invent machinery to synchronize them 7) prematurely generalizing one-off flows 8) timeouts on everything 9) production code that exists purely to satisfy tests 10) patch bad premises additively instead of stepping back and deleting them
a form of slop i'm not seeing on the TL yet is AI's tendencies to: 1) over-engineer solutions 2) overly defensive programming 3) hyper-fixate on rare/fictitious edge cases these tendencies can be clearly traced to how RL post-training is currently done. we need solutions here
20
19
356
21,100
arxiv.org/abs/2608.18066 finds that self-improving agents amplify run-to-run variance and can flip from gains to losses when task order is shuffled. It also shows how unsupported actions, evaluator bugs, and rewarded workarounds produce misleading memories that propagate into future. Questions to authors: Could fixed-memory replay and crossed interventions show whether failures originate in construction, retrieval, excessive authority, or missing validation—and how generally they apply? How much improvement can be squeezed from fixing the bugs in the evaluator code that you found? @qinyuan_ye @yooli23 @yadapruksachatk @jxzhangjhu @jasonwu0731.
1
1
4
294