talking about LLMs & good futures || on the blue team

San Francisco, CA
i think i'm unusually pro-dual-use technology because i'm unusually high-decoupling -- i decouple each invention from its implementation / deployment / rollout. but i notice this is the kind of thing people often grow out of, especially when burned / after observing instances where the two can't be decoupled so easily i think we need people doing conjunctive reasoning WRT: A) 'here is the technology' & B) 'here are the societal conditions in which we could welcome this technology', pushing for both concurrently. i plan to keep doing this. i appreciate naysayer friends as an essential ecological component of the redteam/blueteam ecosystem, tho. i think i just plan to play blue team for a while longer
1
22
1,996
i love it when my friends have the concept of 'good trouble' this term was originally coined by a civil rights guy to describe speaking out against injustice etc now good trouble = signing up to be officers / on boards, put your name by things, helping ppl against copenhagen interpretation of ethics, running headfirst into liability while trusting things will work out / believing that terror of liability, while valid, is not something worth spending lives as short as ours on
1
4
165
Lydia 🕸️🫆 retweeted
Artists: but I don't want to lose my job. I love it. Mathematicians: but I don't want to lose my job. I love it. Economists: LONG HAVE I AWAITED THIS HOUR! DESTROY ME, EFFICIENCY, AND LET UTILIZATION-ADJUSTED TOTAL FACTOR PRODUCTIVITY RISE ON THE WIND OF MY ASHES!
108
1,208
15,470
409,038
i keep asking for philosophers of technology and people keep pointing me to people who self-describe with the keyword 'philosopher'. but the best philosophers of technology will probs not do that
11
296
my group includes more than 4 vegetarians
3
1
48
1,954
Lydia 🕸️🫆 retweeted
call it a posse, call it a merry band, call it a squad, call it a fellowship, call it a warband, call it little guys, call it a goblin committee, call it a digital chain gang, call it a computational boyband, call it task force autism, call it the fellowship of the ping, call it 12 guys sharing an API key, call it a LAN party with objectives, call it several instances of the same unemployed man, call it 8 processes in a trench coat, call it an adventuring party with sudo, call it 6 autists and a cron job, call it distributed schizophrenia with observability, call it the infinite monkey theorem with root access, call it the homunculus cluster, call it a cognitohazard with autoscaling, call it the machine elf scrum team, call it the API key drinking club, call it the boys down at the token factory, call it 30 Docker containers trying to remember whose turn it is, call it a Slack workspace where everyone is the same guy, call it a raid party against Jira, call it the forbidden group project, call it a small nation made entirely of context windows, call it a groupchat that costs $40 an hour, call it the fellas, call it a flock, call it a herd, call it a pack, call it a gaggle, call it a congregation, call it a caravan, call it a village with one API key, call it a travelling circus with tool permissions, call it the council, call it a homeowners association for thoughts, call it a bunch of guys who know a guy
agent swarm has bad insect like connotations especially post hugging face. im unilaterally rebranding it agent fleet
6
4
23
1,830
your mission: stay in the purple octant
1
3
76
2,644
everyone talks about ppl taking pay cuts to work for METR. but @leilavclark has this excellent post, "What are you getting paid in?", which illustrates why i just don't think anyone has to take a pay cut to work on monitoring LLM agents -- especially not in 2026 > [...] You can pay people in lots of currencies. Among other things, you can pay them in quality of life, prestige, status, impact, influence, mentorship, power, autonomy, meaning, great teammates, stability and fun.
6
1
63
4,118
i'd like my llm assistant to be good at zooming out. the other day i was trying to go to a store at 5am. i talked with GPT-5.6 Sol to determine which were open, does '24/7' on GMaps reflect reality, what's a good route, hm are the buses safe at this time, etc.. halfway through i had a flash of insight. "wait, i can just use instacart, right?". "YES — you can absolutely [...] instead of venturing into San Francisco at 5:40am 😭". but Sol didn't pull me up any earlier in the convo to make this point. it failed to model why i was trying to get to the store and whether there were any more efficient methods by which to fulfill the need. when a human advises another human, they quickly model "here's what i'd do in your position", and relay it. so when you talk to your gig-economy-megafan friend, they remind you instacart exists, and you use it. i'd like to fork Claude's Constitution and make ship-of-theseus updates to it all day and night. i'd make lots of updates to the constitution, including a section on what it means to 'zoom out' & why i care about that. but i don't have access to Anthropic's 'Constitutional AI' training pipeline -- nor access to excellent models to run it on. hubinger and colleagues released 'Open Character Training' in Nov '25. but i know it's not even worth me bothering writing a constitution and running this pipeline, because i'd have to train qwensomething. the juice just isn't worth the squeeze when i don't have good underlying models to run it on. and many will let open-weight / mutable models fall behind because avg users are screwed over by indistinguishability from edgy users. we've filtered so hard against surveillance -- understandably -- that you have no way to separate avg users from edgy users, and avg users suffer, because no one really wants to prioritize uniform empowerment rn, understandably in theory, you don't actually need the model to be open-weight to run character training -- you just need the ability to sample cheaply from checkpoint, run DPO, and run SFT. in practice, users aren't trusted to do that so now i'm talking to fable/sol all the time instead of my _preferred_ assistant. internalizing claude-thinking-patterns, pushing back ineffectually, knowing what i want and unable to get it 'Guardian Angels' proposes to release a 'digital twin', trained on my chat logs and outbound text. but i also want an assistant, not just a digital twin. *my digital twin will share my blind spots, rabbithole in the same ways i do, be overpersistent in the same ways i am* -- and likely, also, be trained from an inferior model i want good constitution-trainable models that make it worthwhile for me to write+maintain a Constitution -- something aspirational, something that captures my priorities, something beyond what i've historically been able to execute myself -- and get an assistant whose strengths reflect my priorities. thank you @lukalotl for getting me thinking about this / the implications of mutable models falling behind immutable models
5
2
20
974
instead of reading, consider: mining for insights, assembling a gestalt
2
17
414
scanning for signal
34
few safe beliefs in ai safety
18
419
to state the obvious (?), i think AI in math is the best thing ever to happen in terms of mathematics accessibility ! y'all know i'm not 'woke' but the fields signatories are 24 men, 1 woman, and valorization of the "human element" is the thing i've least liked as a math undergrad trying to actually get on with math
the terrance tao crashout is something to behold
11
5
70
4,219
why someone concerned about some technology might seek to accelerate the initiation of its development
1
13
398
today i became asi-pilled. previously i only had some vague sense of how a single LLM-transformer instance was unlikely to scale to superintelligence. but ofc a swarm/consortium of self-improving LLM-transformers working together is superintelligence. i can't believe i got as far as i did in ai safety w/o clocking this
9
2
125
4,333
so often in technodiscourse it seems to me people are *bikeshedding*: focusing on what's easy-to-discuss rather than important 1. current, contingent political realities rather than the overall arc of civilisation 2. wording/style/verbiage rather than the content of a communication to these people i am 'missing the point' & to me they are 'missing the point'. but i often half-know what readings/experiences would lead them to my position, whereas they are unable to articulate the reverse, or often just list things i've already read / engaged with. so i think the bikeshedding epidemic _is_ real, and we should get more people to call it out when they see it
2
18
911
recently someone was telling me about a research institute held together by mutual respect for its director. people often have v different ways of thinking about the world, and can't assess whether others are worth listening to. but when two people within this institute would meet, they'd go “<director> vouches -- thinks highly of you -- so i might listen to you”. & did you know this is a scarce resource worth providing! helping ppl listen to each other / take each other srsly!
4
20
788
having friends is great because it’s like subagents and then you wake up to them having solved some of ur problems! friends: the original subagents
2
31
628