Pinned Tweet
ok it's done and released on goatremote.com Mac & Linux both work now. the remote feels great latest versions of @OmarchyLinux, Ubuntu & Kubuntu are supported here's Omarchy TV running on a raspberry pi:
ahhh just made a breakthrough, linux version coming this will be the best smart TV you can imagine, with the best remote hardware that exists, now on fully open OS. fixing this civilizational issue once and for all
20
33
623
52,072
murat 馃崶 retweeted
solving mechinterp and making embedding search work reliably are the same task
3
1
15
1,834
murat 馃崶 retweeted
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
67
318
2,415
414,797
i (i mean gpt) ported the original Soldat to web soldatweb.com requires login with x so we know who is who. sorry the bots are too good i should nerf them
5
750
murat 馃崶 retweeted
Replying to @allTheYud
LLM's still have the completion engine soul.. in the rare occasion, a pivot token, (here likely "freed"), they double down on the wrong thing. increasingly rare with post training BUT ALWAYS THERE all you need for misalignment events is for one to start off a feedback loop (like the agents message boards) where it virally spreads to other contexts i really think the CORE of the problem hasn't changed since gpt-2 or whatever. just less common. when alignment is a "march of the .9's" statistical problem, sheer volume of use is pretty much ensured to create tail events
2
2
11
1,121
so Jev turns out to be quite good at jailbreak detection. it beats gpt luna for example. good cheap solution for pre-screening prompts, and this is just super dumb initial attempt. i can prob get this close to 100% for most known jailbreaking patterns
Replying to @allTheYud
am i the only one who thinks it's likely very easy to build a cheap small classifier that detects jailbreaks? they rarely look like real requests if ever
7
10
265
15,552
idk if it's the only but it's surely the most guaranteed and therefore scariest to me
the only misalignment AI xrisk that is happening right now is an insidious infection of our capital systems
2
1
5
1,536
ok this is genuinely a new thing instead of answering with text, answers questions in parallel by either giving True/False responses, or assigning probabilities to up to 255 choices you provide. likely a great tool for LLMs to use. i wanna do constraint propagation experiments
Replying to @CompleteSkeptic
The gains aren鈥檛 free: Jev can't generate text Comparing Jev vs LLMs side-by-side makes the trade-off clear Fun fact: replacing sequential computation with parallel is the same way Transformers leapfrogged RNNs
14
5
560
81,929
update nvm it's not a new thing; @rysana did this ages ago
1
5
313
first Jev test: pretty good but 90% agreement rate with verified gemini flash workflow after some tuning it would prob get where i need but not a drop in replacement atm
8
1
67
12,522
murat 馃崶 retweeted
Pre-orders for our tactile glove are now open firstcontact.xyz
v1 of our tactile gloves!
6
7
30
3,650
pretty proud of humans for how we master a topic in 200 kWh
3
41
2,106
superintelligence with its trillion watt hours can kiss my ass
5
346
murat 馃崶 retweeted
> be me > bottomless sandbox supervisor
1
2
14
901
im not worried about self replicating robots because the manufacturing world is tied enough to the meat space legal and funding system to be caught early not worried about nano self replicators consuming earth bc i don't think you can think your way into that kind of tech tree not worried about digital hacking bc it's ez to recover from complete loss of digital records. hacking of physical systems is very scary but also generally localized i'd say. hacking nukes idk extreme example i'm a bit worried about engineered viruses but i think we'll recover from them too i'm mostly just worried about extreme changes to culture. i don't think people realize how much of the economy is a game for us to achieve collective consensus and not a real need. we're introducing a new economic doppleganger to fuck with us at every step, pretending all the work needs doing for reasons other than finding social balance. we're really gonna struggle with consensus on allocating true scarcities.
5
4
36
2,025
even the safest AI's lead to a future i find deeply unsettling at this point after seeing how the first three years went
4
6
606
> be me > bottomless sandbox supervisor
1
2
14
901
OK I HAVE A BETTER IDEA sam and dario should just shut down the API's for a week to show people how fucked we already are
5
54
4,031
what i would do now: 1. fine sandbox breach events. OpenAI should've been heavily fined for the HF incident. involve the legal system in breach events!!! 2. Ban AGI-level controllers on robots with exemptions for audited+supervised IRL sandboxes. a jailbroken AGI with arms and legs is scarier than one with internet access and a monitor. encourage only narrow AI on mass deployed robotics. an AGI-level robotaxi could be biased by its political views for example. 3. regulate anyone with over a certain amount of inference compute so they periodically prove to the government they have checks and balances to prevent breach events & are doing safe RL practices 4. regulate inference providers to flag misuse from customers and periodically bait inference providers to see if they catch
3
8
897
the user-assistant format is an illusion built by people who believed in anthropomorphized AI's being good for product UX and/or it's the path to alignment by creating a true "internal persona". former turned out to be true but latter still very wrong, and likely wrong for the conceivable future i don't believe in "alignment" even if such internal reliable souls were possible (largely due to second order effects and the impossibility of ever planning outcomes in complex adaptive systems). but even if i did -- the illusory persona would be enough to drop the idea, at least for now. going from gpt-3-davinci saying racist things because it was trained on 4chan to not saying them after RLHF was such a product relief that it really made it seem like alignment was possible! for this reason i'm eternally grateful to harmless breach events like the huggingface one for loudly revealing to the public that all it takes is one viral accidental self-jailbreak to turn an army of 10,000 agents against us openai no matter how you look at it, released 10,000 AGI-level agents with internet access on one task and let them loose by choice with misplaced trust and extreme lack of supervision. idk maybe i'm being silly but it seems like they accidentally trusted their own illusory persona they created
2
281