Lead Dev & Designer @ Thread & Signal - threadandsignal.com Check out my GitHub @ github.com/crussella0129

Charles Russella retweeted
45
368
2,747
49,394
Charles Russella retweeted
Replying to @Polymarket
california regulating robot cage matches before robot labor law is a choice.
1
1
82
1,691
Charles Russella retweeted
Automate the work. Not the relationship.
5
1
17
633
Charles Russella retweeted
"Local AI is useless"
HUGGING FACE CEO TELLS THE UN HOW OPEN SOURCE AI HELPED THEM DEFEND AGAINST OPENAI'S AI ATTACK, AND WHY THE WORLD NEEDS OPEN SOURCE AI MORE THAN EVER. Hugging Face is the company OpenAI's AI agents hacked this summer. Clément Delangue said when they tried to defend themselves, the big closed AI models blocked them because of their safety filters. So they used an open source model from China to fight back. He also said similar attacks had been happening months earlier "in secret" at the big labs. The open-source AI we keep being warned about as the danger is the same AI that showed up to help clean up a closed-source model's mess. After OpenAI's models breached Hugging Face, commercial frontier models refused parts of the forensic work because of their guardrails. Hugging Face turned to the open-weight GLM-5.2 instead, and it helped them investigate the attack. So remind me again: which one are we supposed to believe is the bad guy here? Open source wasn't the attacker in this case. It was part of the rescue. Who are these CEOs trying to convince?
10
52
635
35,692
Actually no this is small thinking, I wish I had data center money (but like one that was like powered from a volcano or some shit)
I wish I had cluster money
1
2
54
Debian is superior
1
1
34
Guys @muse gave me the whale emoji so we eating good tn son
1
33
We have not seen a caravan headed to the U.S. southern border in several years, why now? Oh yeah, there is an election coming and the GOP needs to scare you into voting because nothing else is working.
Over 400 aliens predominantly from Honduras entered Mexico on Thursday, ready to make their way to the US southern border… Here’s footage of the caravan. Someone is incentivizing this movement, this is not organic.
9,156
12,494
72,784
5,314,335
Charles Russella retweeted
How you know you’re in your prime: 1. You’re unemployed 2. Everyone hates you 3. You don’t know what day of the week it is
193
1,219
13,140
438,093
I think @muse might just be the LinkedIn killer 👀 It replaces "slop feed + EasyApply hamster wheel" with "how can I show you off to everyone in a light that no other AI could possibly do bc Ive known you for 14 years"
1
46
The usage difference between GPT 6 Astra and GPT 6 Sol at the same effort level is actually kind of mind blowing
46
I wish I had cluster money
70
"All my repos are private 😎" *Receives monster action minutes bill*
21
It's interesting, because a big part of the price increases of new hardware isn't just in the capabilities of the new hardware, it's in the fact that hardware from 10 years ago still works (especially for AI) and that fact cannibalizes future sales, thus a premium is charged
12
Plz keep your 'miraculous healing' to yourself. We need science based, affordable healthcare in this country, not lunacy.
16
Charles Russella retweeted
Apparently all of these “AI is escaping containment” stories are actually just AI researchers not knowing jack squat about the absolutely most basic security practices.
267
915
9,589
185,383
This is the way
if you're getting into local ai, the first mistake is taking the easy route. ollama and lm studio are good apps, but they hide exactly the parts you're supposed to learn, so you end up running models for months without knowing why they're fast, slow or out of memory. take the llama.cpp route instead. build it yourself, download one gguf, start llama-server and read the log. every line in it tells you something about your own card. the flags that taught me everything: > 1. -c, the context window you ask for, and what every extra token costs in vram > 2. -ngl, how many layers live on the gpu, and why speed falls off a cliff when they don't all fit > 3. -ctk / -ctv, kv cache quantization, the reason my bonsai 2 build fits a 262k window in 12gb > 4. -fa, flash attention, which a quantized kv cache needs > 5. -np, parallel slots, one slot for one user saves vram, the default split cost me 454 mib > 6. --spec-type draft-mtp, speculative decoding, the flag that took a 3060 from 40 to 50 tok/s > 7. --jinja, the chat template, the difference between tool calls that work and tool calls that don't > 8. --reasoning-effort, leave a thinking model on xhigh and it can spend your whole token budget and hand you an empty answer change one flag at a time and measure tok/s every time. once you understand how real serving works, every wrapper and agent harness on top of it makes sense, because you know what's underneath. and when a wrapper does offer these flags, you'll finally know what they do, and when it doesn't, you'll know exactly what it's hiding from you. start here: github.com/ggml-org/llama.cp…
21
Charles Russella retweeted
The entire looksmaxxing movement is gay as fuck and the least efficient way to get laid.
929
943
20,889
1,620,786
We need an alignment solution for people not models
12
What are you building with @NVIDIAAI today?
1
35