Filter
Exclude
Time range
-
Minimum likes
Replying to @0xSero
This is the way. 2 of these blow a Spark out of the water in terms of performance.
208
Alex Brandes ² retweeted
There's nothing kind about letting good people live in a fantasy world that no longer exists. You have to tell them, even if it hurts. Because the sooner they accept reality, the sooner they can adapt to the future.
260
557
7,600
691,615
Replying to @joshEclass
Absolutely no complaints. I’ve pushed billions of tokens through already. Qwen 27b at 8x concurrency or Qwen Flash Next at 6x seems to be the sweet spot.
1
1
22
Replying to @joshEclass
eBay. Prices keep climbing though.
1
40
This is an absolutely wild price for 32GB VRAM.
RTX 5090s are now $7,500 at Best Buy
2
4
479
I finally got my CMP 170HX LLM server running, and this thing is an absolute beast: 128GB of VRAM across two cards, running Qwen3.8 27B at 8x concurrency. Single-stream decode is already above 150 tok/s with DFlash, with aggregate throughput approaching 200 tok/s. I’ve barely optimized the workload. At API prices, it could easily burn through $100 a day. Working on some interesting experiments with agent swarms to soak up all that bandwidth...
6
1
15
726
Replying to @alexocheema
Sure. @antirez built DS4/DwarfStar specifically around DeepSeek V4 Flash instead of waiting for a general runtime to absorb every model-specific optimization. @Youssofal_ MTPLX brought native MTP speculative decoding to Apple Silicon when MLX/GGUF/LM Studio didn’t have it. @ddalcu’s mlx-serve is now showing things like +154% from native MTP on Qwen 3.6 35B and ~2x vs LM Studio on benchmarks. These are all beating the “general purpose stacks” in terms of performance. Specialization lets people aggressively exploit a model’s architecture and a hardware target instead of being constrained by what makes sense for the entire ecosystem. This is becoming significantly more optimal with LLMs. Let’s support these contributors.
3
2
28
518
Replying to @alexocheema
I absolutely hate this take. Specialized inference engines have driven insane performance gains in just the last few months, pushing what’s possible on specific models far beyond general-purpose stacks. Calling that “slop” and ecosystem fragmentation feels completely at odds with what your project is trying to accomplish.
1
5
1,631
Replying to @ItsmeAjayKV
CMP 170HX + Qwen3.8 27b - ha!
1
1
326
This is the truth and I think the people that are deep in it also see that we’ve only barely scratched the surface. I’d rather be in the trenches than sitting on the sidelines.
Anyone that works in AI is working the hardest they've ever worked in their lives. On the surface it's somewhat ironic (AI should give us back time!), but the reality is that it's the most fun, fascinating and empowering epoch in human history. The intelligence revolution.
5
518
You can do this with tmux on its own, or with wrappers. I find this works really well with Orca which gives you the option to see tabs of all the sessions or run the in the background if you want. Agents coordinating agents becomes a necessity at a certain point. Subagents in a single session is not the same thing.
I’m seeing lots of tweet replies asking basically “but how do I get two models to talk to each other? Without me manually copy/pasting results between apps?” And I’ve heard this same question repeatedly from friends who use AIs mainly through GUI apps. Yes, there are various apps starting to be released for just this problem. But it’s worth knowing that in a pinch you can already do all this yourself, on the command line, just by using tmux. Here’s how. On the command line start a named tmux session with the command “tmux new -s chat”. Create a second window within that session. (The default keystroke for this is doing “ctrl-b c”. You can switch between windows by doing “ctrl-b n” for the next window.) Launch the codex CLI in one window, and the claude CLI in the other. Then you can literally say to the AI in window 1, “please read the analysis of the AI in window 2 of the tmux session ‘chat’. Wdyt?” And vice versa. Now you have cross-harness review. Or you can instruct one AI to delegate and manage the other AI over multiple turns, or monitor it over time, or whatever. This works because each AI can use the tmux CLI to read and write to the other’s window. This does not on its own solve coordination problems like deciding who is in charge, or who talks first, or who waits for whom, but it establishes the basic communication primitives of “read the other AI’s output” and “write to the other AI”. You can build more complex workflows on top of that mostly by prompting. The fundamental reason this all works is that the Unix-era command line environment is more interoperable than the modern GUI environment, so a decades-old tool like tmux is still one of the best ways to compose modern AIs. Obviously the command line tools are a PITA in other ways. But it’s great that they are so flexible, they already exist, and they work right now, so you don’t need to wait for the vendors to get their apps to play nicely with each other (which might never happen) or for someone else to catch up and write a new app just to connect things.
1
5
417
👀
If you ever think you can’t do something remember moltbook got acquired
6
431
Oh wow.
the openai huggingface incident, from an agents pov. (part 1)
7
450
Replying to @nateberkopec
I get 20% extra tokens by using it through @_GateAI but also check your effort level. Low effort level is better than Opus 5 high and doesn’t burn through the limit too quickly. Do not let it launch subagents! lol
1
3
167
Astra is looking really impressive across pretty much every benchmark. Kind of wild that turning up the reasoning actually makes it cheaper. It takes fewer steps to finish the task, so the total cost goes down.
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation. Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself. We see Astra as a major breakthrough in model intelligence. Read our post on Astra and what these results mean: arcprize.org/blog/astra
1
6
385
A lot of people are sleeping on how big a deal it is to generate video faster than it can be played back. Fast H3 can generate a video before you finish watching it, opening up a whole new category of real-time video applications. This is the beginning of video becoming an interactive interface, not just content.
🎬Video generation faster than playback! 🚀MiniMax H3 on vLLM-Omni + FastVideo's FastH3: a complete 10.1s MP4 - video AND synchronized audio - rendered in 8.7s!⚡️ Thanks to @MiniMax_AI for the great Minimax H3 release, the FastVideo team @haoailab for open-sourcing FastH3 and helping on the serving integration, and @NVIDIAAI for the continued sponsorship and joint optimization efforts!
1
2
13
606
The defects of the past are the raw material of the present. Let’s get weird!
I think the CMP 170HX just became one of the weirdest local-LLM value plays. 👀 I got Qwen3.8-27B W8A16 running on a single 170HX 64GB with vLLM 0.27.1 + DFlash2: ⚡ ~120 tok/s decode 🚀 ~1.6–1.9K tok/s prefill 🧠 262K context with BF16 KV 📦 342K-token KV capacity 🎯 DFlash2 averaging 3.66 tokens/draft 💾 60.5 / 63.5 GiB VRAM used 27B model is doing ~120 tokens/sec on one GPU. Software optimization is getting ridiculous. Those numbers are from your measured W8A16 + DFlash2 run: ~120 tok/s decode, ~1,580–1,940 tok/s prefill, 342,729-token KV capacity, and 60.5/63.5 GiB VRAM usage. #170hx #LocalAI #localllm
4
448
Replying to @alexocheema
The ever-present desire for more compute.
9
397
Fable 5.1 low scores higher than Fable 5 max and is much cheaper. No need for Opus anymore.
Replying to @claudeai
As well as being capable of much higher performance than Fable 5, it can also achieve similar or better results at a much lower cost when set to lower effort levels.
3
451
Replying to @benhylak
Yes
People still underestimate what it means that software can write itself. You don’t need to become a programmer. You need an agent like GPT or Claude, maybe some cheap hardware, and to describe what you want. 15 things you can build now with zero programming knowledge: • Create your own desktop photo app that organizes pictures exactly how you want, including custom tags, face grouping, duplicate cleanup, and private local search • Build a private, local Google Home replacement with a touchscreen, microphones, speakers, cameras, and voice control, with nothing sent to the cloud • Create a 3D model of your house and yard from satellite imagery, drone photos, and measurements so you can test trees, patios, fences, lighting, and landscaping before spending money • Buy a cheap ESP32 board and build your own Bluetooth hardware, like a custom remote, sensor, button panel, bike computer, or desk controller • Make a custom dashboard for an older car that reads OBD-II data and displays exactly the gauges, alerts, maintenance info, and trip stats you care about • Turn an old tablet into a purpose-built kitchen computer with family calendars, recipes, timers, grocery lists, music controls, meal planning, and a UI designed for your house • Build a local security-camera system that can tell the difference between “a person is at the door,” “the dog wants to come inside,” and “a package was left by the garage” • Make a private family-photo archive that identifies people and places, reconstructs old trips, groups photos by event, and creates a searchable timeline across decades • Build a computer-vision system for your garage or workshop that recognizes your tools and can answer “where did I leave the 10mm socket?” - Put a camera in your garden that tracks plant growth, identifies weeds or pests, watches for animals, and shows you exactly what changed overnight • Build your own baby monitor that runs locally and tracks crying, movement, room temperature, sleep patterns, and whatever alerts you actually want • Make a custom e-reader that can explain references, track characters, build timelines, translate passages, summarize footnotes, and adapt itself to the way you read • Build a completely local “second brain” that indexes years of PDFs, screenshots, notes, recordings, receipts, manuals, and documents and lets you ask questions across all of it • Create a radically simplified computer interface for an elderly parent that only exposes the five things they actually use, with huge controls and custom workflows • Build a controller for a weird thing in your life: your aquarium, greenhouse, chicken coop, telescope, camper van, 3D printer farm, wine cellar, workshop, or home energy system If you can describe what you want, you can probably build it right now. Also, try promoting by voice! Your imprecise ramblings actually add context that produces a better end result.
2
108