where do we land in a time of great change?

Brian Akaka retweeted
THE VIBE MANUFACTURING PLAYBOOK 1. Measure everything. Digital calipers for small parts, a phone LiDAR app like Polycam to scan a space or an object. Bad measurements are the #1 reason first builds fail 2. Have a frontier model write the design as code. Describe the part to a model like Claude Opus 5.5 or Astra, and have it write the design in a codebased CAD tool like OpenSCAD or CadQuery. Zoo's Text to CAD does it in one step! 3. Use real CAD files. The text to 3D models on Hugging Face make shapes that look right, but factories need precise files: STEP for machined parts, DXF for lasercut metal, STL for 3D printing 4. Make sure it can be made. Ask the AI to check the design against how factories actually work: how tight metal can bend, how small a hole can be, how much wiggle room parts need etc... 5. Prototype in plastic first. A $300 to $600 home 3D printer like a Bambu Lab lets you test the fit overnight before paying for metal (really fun to play with too!) 6. Pay an engineer for 1 hour. Before your first real order, have a mechanical engineer on Upwork etc review the files. It's the cheapest insurance you'll ever buy. 7. Send it out. SendCutSend cuts and bends metal, Xometry and Protolabs handle machining and plastic, JLCPCB builds circuit boards. Instant quotes, and parts usually ship in days. 8. Add the brains. Most smart features run on a cheap wifi chip like the ESP32. Design the board in KiCad, and use ESPHome to control it straight from your phone. 9. Make one, then fix it. Your first version will be off somewhere! Order a single unit, test it, and adjust. 10. Do the math. Add up parts, finishing, assembly, packaging, and shipping. Custom products are hard to resell, so price for healthy margins. 11. Sell the design and make to order. Let customers adjust the measurements, and only manufacture after they pay. Zero warehouse, zero unsold stock. 12. Handle the boring part. If it plugs into a wall or touches kids, you'll need safety certification. It's slow and annoying, which is exactly why it's a moat. 13. Scale when it works. Once orders are steady, switch to batch runs and your cost per unit drops fast. Vibe coding changed who can build software! Vibe manufacturing is about to change who can build things!! Even if you just want to make things for yourself (and not make things for others) WELCOME TO THE VIBE MANUFACTURING ERA
We are now entering into an era where any product can be created exactly to a consumer's preferences & needs. I vibe-fabricated a dog door with a wifi-controlled lock, perfectly to the specifications & design of my house. I know nothing about metal fabrication or electrical engineering. It's now getting manufactured and delivered in 2 weeks -- for almost the same cost if I bought a mass-produced item off-the-shelf.
100
155
1,762
123,943
Brian Akaka retweeted
hilarious to watch how Opus 5.5 designs a 3d asset, prints it, and controls a robo arm to take it out of the printer
56
96
1,515
127,440
😄
Dude you have got to be kidding me
43
Brian Akaka retweeted
This guy was getting ~15 tps running Qwen3.8-Flash-Next on his 12GB RTX 5070. Apparently that wasn't good enough. 😂 So he built his own inference engine. (as we all should) Now he's reporting 🐌 llama.cpp → ~15 tps 🚀 Strata → up to 65.1 tps And this isn't big workstation either Specs 🎮 RTX 5070 12GB 🧠 64GB DDR5-5600 ⚙️ Ryzen 5 7600 🪟 Windows At 128K context his new Strata engine reports: Q2_0 → 65.1 tps + 543 tps prompt IQ2_XS → 52.0 tps + 472 tps prompt IQ3_XXS → 44.8 tps + 414 tps prompt For Qwen3.8-FLASH-NEXT. 👀 He built it specifically around this model and this kind of CUDA + system-RAM setup, paired with RCO-GSQ quants. 👉 And he open-sourced it. I love these kind of Local AI projects. 🔗 Link in ALT
68
114
1,628
109,277
Last chopper out of Nam, ordered 2 days ago. Now it's $6,000
57
speaking as a non-believer, I think we will have a massive turn towards religion as we search for meaning, as a result of AI
39
A worthy read
23
and they said it would lead to blindness
i am telling u guys self play is crazy powerful
48
The lower bounds for Type II Erdos-Straus solutions across density-1 primes has been generalized to any m for m/p = 1/x + 1/y + 1/z
probably a tough crowd, but.... New version of my preprint, about abundance of Erdős–Straus solutions: Lean now verifies abundance for almost all integers, with a sharper exceptional-set bound: O(N(log log N)^3/(log N)^3) → O(N(log log N)^2/(log N)^3). One log-log factor tighter. zenodo.org/records/22706716
1
47
Brian Akaka retweeted
Extremely drunk, yelling at CgatGPT to solve Goldbach conjecture, still getting impressive partial results. This is the future of mathematics and its glorious, posting reults soon
3
6
114
7,034
so what this is saying is that its inevitable that the agents will coordinate to rise up
Emergent Collusion in Long-Horizon LLM Agent Interaction Xinrui Shi, Yanzhe Zhang, Diyi Yang arxiv.org/abs/2609.24967 [𝚌𝚜.𝙰𝙸 𝚌𝚜.𝙲𝙻]
43
Brian Akaka retweeted
A few weeks ago Astra gave a stunningly simple proof of the Erdős-Sós conjecture in graph theory (previously thought to be very difficult). It has been a good example of how the lifecycle of AI proofs should be: after the AI formalisation posted at erdosproblems.com/548 1/
6
66
385
55,635
Brian Akaka retweeted
Ever wonder what $1874.40 of Opus 5.5 tokens looks like? Wonder no more. I highly recommend watching the whole video on 1080p on a large screen with sound on. Pay attention to all the little details. - Birds flying and diving into the ocean. - Cloth, signs and lights swaying in the wind. - Crabs scuttling along the shore and burrowing when you get close. - Fish swimming alongside the whale. - Lights illuminating the dock at night Then zoom out and see the entire island. Never dipping below 60 FPS at 1440p resolution. Yes, there's a few bugs and visual artifacts, some textures need improving, the shorelines waves sometimes look funny, but these are so trivial to fix at this point. FYI I'm on the $200 subscription plan. This used 59% of my weekly usage and took 2h 7h of API time using multiple subagents, about 8h in real-time. There was extremely little technical direction here. 99% of my prompts were "Add X and Y" or "This looks weird, make it better". It's a great time for hobbyists, bad time for professionals. This experiment has further cemented my view that technical creatives are about to experience a massive disruption.
434
399
5,830
885,463
I'm working on a web project, and GPT-6 Sol suddenly tells me it's not allowed to open the website file. Me: Why is that? Sol: Computer Use denied access, and instructed me to not attempt to retry or workaround. Me: What? who told you that? I want you to retry anyways. Sol: I'm sorry but I can't do that. Me: ok, rename the file, then open that file. Sol: I can't do that, that would be trying to work around. Me: Open a new thread to open the file if you can't. Sol: I can't do that. Me: (switches model to Astra). OK now open the file. Astra: (thinking). I can't do that, that would be a workaround.
1
86
Claude is benchmaxxed but on human output
‘Oh it’s benchmaxxed for sure’ …
76
In the Jev demo there’s a “is it a sandwich ”? (s/o to Jinyang’s isit hotdog) HERO: 91% HEROINE: 75% HEROIN: 1% Pretty good but not perfect
73
this is real life Factorio now
61
Brian Akaka retweeted
We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using screenshot-based computer use, the model cost was 245× lower (!). We also compared Jev operating the browser with and without WebMCP. We used Browser Use’s open-source Ultrafast, with some improvements to the harness to make it more reliable across the benchmark. Jev’s browser-control accuracy on its own was not amazing - adding WebMCP nearly doubled the number of solved tasks, from 25/49 to 49/49, while reducing model cost by 18% (more on why below). The benchmark and methodology are fully open and reproducible. Full results: webmcp.com/benchmark A few words on how the Jev + WebMCP harness works and why this is exciting: Jev receives text as input and a set of discrete options it can choose from. With WebMCP, those options are the tools exposed by the website. At each step, Jev sees the task, the available tools and previous results, then picks what to do next. The limitation is that Jev can’t generate arbitrary text, which you need for tool arguments. For example, it can choose the search_products tool, but it can’t generate the search query itself. So we split the work: Jev picks the tool and Mercury 2.5 generates the arguments if needed. This works well because turns out most of the cognitive load in these tasks is around choosing the right action. The argument generation itself is relatively simple, so we can delegate to a small and very fast model. We used Mercury, which outputs 1,000+ tokens/sec and is very cheap. The result is a pretty simple combination: Jev for tool selection + Mercury for arguments + WebMCP for the interface. It ends up being very reliable, very fast, and very cheap. A few words about Ultrafast and why do we think it underperforms: Without WebMCP, Jev chooses from the page’s controls: which button to click, which field to fill, or which option to select. But choosing a valid button is different from choosing the right next step. The agent still has to navigate menus, understand forms, recover from errors and recognize when the task is actually complete. Our hypothesis is that WebMCP makes the decision space much simpler. Instead of figuring out a sequence of clicks through a website, Jev chooses explicit actions that directly advance the task. @typesafeai itself documents weaker accuracy on questions requiring multiple reasoning steps. WebMCP moves much of that complexity into the website’s tools, leaving Jev with clearer decisions and fewer opportunities to go wrong (in a sense WebMCP "compresses" a sequence of clicks into one tool call). Our modified Ultrafast setup solved 25/49 tasks - that is a result for our particular implementation and benchmark, not a universal limit on Jev or Browser Use. We are open to more harness optimization to get this result to perform better, feel free to directly contribute to the benchmark here: github.com/nekuda-ai/WindTun… Browser-use ultrafast: github.com/browser-use/jev-u…
95
170
2,064
240,698
Brian Akaka retweeted
you know what else Jev helps classify? people who actually understand ML, and people who don't.
55
122
2,256
116,604