I saw a video today with a chart that was a bit confusing at first, but then it got. It perfectly proves hardware determines your baseline, but your software framework dictates your absolute ceiling. Look at the crossover point. At 1 to 4 users, llama.cpp thrives on raw local bandwidth. But at 8 users, it hits a context bottleneck. Meanwhile, vLLM's continuous batching and PagedAttention keep scaling right through it. Secret sauce is PagedAttention if you wanna read. Beautiful architecture is always the thing I appreciate the most, and vLLM is exactly that.
1
Okay, working on a terminal project in Rust. I'm keeping optimization as a top priority It’s pretty rough right now. I'm just working on the thing and learning as much as I can. If it turns out great, maybe I'll make it for other people to use down the road. I’ll be sharing updates whenever I get a moment. At the end of the day, it's all about the experience and I'm just passionate to do it. Getting rusted up!
11
Finally X has removed the ghost ban and I am seeing my posts and replies as visible
11
Got to know X is ghost banning me unnecessarily even when I'm not posting much wow and X has worst technical support @premium
8
Okay I was not around for a day and google already launched gemini 4. They did not name it pro they named it Argon. I guess they want to make it feel different.
10
I was watching a friend’s workflow take over 30 minutes to complete. After looking into it, I fixed the pipeline and brought it down to roughly 4 minutes. The lesson is this: I believe the problem with most of these getting wrong is when agents themselves figure out CI/CD pipelines and proactively interfere with them. Without proper information or guides, It is causes by issues like: → running heavy tests on every commit → no caching → no parallel execution → especially their automatic setup of older methodologies like ESLint. Agents can write CI/CD configuration.
14
Halving the compute on a $200 plan and calling it a win. The real breakdown: 1. The 50% Compute Cut: Slicing our quota in half while using existing provider price drops as an excuse isn't a benefit, it’s an outright 50% price hike on a $200 plan. 2. The "Perk" Illusion: The no 5-hour limit was already the baseline a month ago, and unannounced "free features" don't compensate for gutting core compute. 3. The Quality Downgrade: The "50% cheaper" models (Sol and Luna) are noticeably worse in real-world output than the flagships we actually pay $200 to use. 4. The Precedent: If this is how you hollow out your top-tier $200 Pro plan, what should users on the lower tiers expect coming?
Hi, Tomorrow we are re-opening the Pro $200 subscriptions to new subscribers, but together with it we are also changing how we calculate the usage for it. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan. Now that it's said, let me explain why this is happening and why you will still get more work done than if you were on the Pro $200 subscription one month ago. (a) We didn't want to compromise in other ways and are committing to not reintroducing the 5h limit, so that you can fully use the weekly usage when you want. (b) On the subscription, we guarantee that over time you always get more work done and with an increasing level of quality. This means that you will continue to get more value per dollar spent as a result of models getting more efficient and us passing down the improvements in the form of API price reductions. (c) We don't want to put an incentive on ourselves to artificially inflate the API list prices to make it look like you are getting a lot (and workaround it through discounts, etc). Instead we want to continue to both rapidly reduce prices and increase capabilities of models on the API. This week we introduced GPT-6 Sol and GPT-6 Luna at 50% of their previous price. Over time, we see prices go low enough that it makes sense for most to buy usage as needed without there being a significant gap between what you get in a subscription and what you get in the API for a dollar spent. (d) Tomorrow, we are adding more things to the subscription that won't draw on the usage, I won't reveal what that is yet. I wanted to be transparent before all the big announcements tomorrow. Lots of new exciting things are coming to the subscriptions that will make it super compelling, but I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news. Codexingly, Tibo
13
fun bug from today a data pipe had a safety valve buffer fills, writer pauses, re-checks every 5 seconds in case things drained. things drained in 3ms. the writer still waited the full 5. my agent after multi-review said latency was solid (which was not the case) the annoying part was nothing was actually wrong. the check was fine, the wait was fine. nobody ever woke the waiter the re-check only ran on the timeout, so every full buffer meant one step per 5s instead of thousands per second. fix was one line. the consumer gives it a nudge when it finishes draining. the nudge carries zero information, the writer still re-checks everything itself. it just needs a reason to look. keeping this one-> a self-checking loop is only as fast as its wake-up source. if a timeout is the only thing that re-runs your check, your timeout is your throughput.
10
It is fake and I'll bet you on that, btw one thing I believe, What is the purpose of 2M? It's really absurd It's the same company which locked all of its models context window to 250k in agy.
10
I noticed something weird on AA recently. deepseek v4.1 flash has a higher verbosity score than glm 5.3 flash while using, it feels like the complete opposite. dsv4.1 feels super brief, while glm loves dumping walls of text. A big part of it comes down to inference architecture and raw decode throughput dsv4.1 uses a causal encoder decoder mechanism with fp4 kv cache quantization and only 16B active parameters on decode. It streams at such absurd speeds that a 1,000-token output flies by in two seconds and registers in your head as concise meanwhile, glm runs an 18B active MoE using pretty much the same hybrid setup Kimi works on (Kimi Delta + MLA). Its per-token latency is higher, and because it’s tuned for long-horizon agent execution, it streams its step-by-step logic and file paths directly into the chat like a verbose system log. Obviously, some of it is just how they measure it on benchmark evals as well. But It is interesting to see how exactly it works.
8
I was working on a desktop app in Rust recently where I needed terminal sessions to keep running even if the user closed the main window. Coming from an Electron and TypeScript background, I almost over-engineered the entire thing. Thought I would share what I ran into. In a lot of Electron dev tools, you will see a separate background daemon process running. It talks to the UI over local sockets so the terminal stays alive if the window reloads or closes. If you come from that ecosystem, it is easy to assume that is just the standard way desktop apps handle persistence, and that you need to build a daemon too. Except that separate process was just a workaround for Electron. When a Chromium window closes, the process dies with it. With Tauri and Rust, that problem does not exist. The webview is just a visual layer. The Rust core process already outlives the window. It owns the OS threads, the PTY sessions, and the memory. When the window closes, nothing actually stops running. When the user reopens the window, the UI just re-attaches to the state that was already sitting in RAM. You do not need a second process, socket code, or reconnection handshakes. You just keep the state in memory behind a clean interface inside your existing engine. Doing it in-process gives you: -Zero socket latency because everything is direct in-memory calls. -About 25MB of RAM instead of 150MB or more for an extra background runtime. -No risk of orphaned zombie processes left behind after crashes.
1
7
What's up with touching grass? I prefer touching keys only.
11
algorithm please #connect me with: - serious builders & indie hackers - based in SF / Bay Area or building globally - ambitious founders obsessed with what's next - saas creators & developers - anyone shipping cool stuff in public say hi or drop what you're working on. let's unite and build.
1
17
Now that the whole Muse mascot wiggling craze is over for now The model underneath just isn’t quite there yet It's a fun companion to mess around with, but it still has a long way to go on actual reasoning.
1
12
I ran 267 benchmark runs across all 89 tasks of Terminal-Bench 2.1 in Claude Code to see how prompt architectures actually behave. Comparing Supreme vs Superpowers by obra vs unprompted baseline: → Supreme: 61/89 solved (20/30 hard tasks) → Superpowers: 58/89 solved (16/30 hard tasks) → Baseline: 55/89 solved (17/30 hard tasks) Supreme solved more complex tasks at a lower cost ($0.489 vs $0.506 per solved task). The difference came down to execution pacing: 42 surgical line edits vs 471 destructive whole-file overwrites, and zero polling traps. dive into the charts, transcripts, and full breakdown in the article below:
1
2
45