every new session i re-explained the same port, the same constraint, the same fix we already ruled out. the agents were never the bottleneck. i was the memory between them. so i built coding brain.
Article

Coding Brain: One memory for every AI coding agent

I run about 40 projects. Most days I have coding sessions open in parallel across Claude Code and Cursor, sometimes Codex too. The agents are good. That was never the problem. The problem was that

20
1
26
8,705
mcp usage across claude products is up 110x this year. the new plugin portal lets developers bundle mcp servers with skills into one package, submit them for review, and track installs by product version. third-party extensions just became easier.
11
docker containerized apps. nix containerized dependencies. personal os containerizes the entire environment. makes sense.
Bullish on open source, open weights, the open web And the open desktop Personal Operating Systems feel like the natural next step after Personal Software I've been building a Debian-based OS for myself that can run across physical machines and cloud environments like Vercel Sandbox I can use and control it with natural language, and I'm designing the UI around a blend of ephemeral and persistent views rather than treating the desktop as a fixed set of apps and windows There's something really special about having control over the UI/UX of the Whole Computer
25
asked opus 5.5 to animate what happens after you hit enter. one prompt, every frame is code.
33
176 test runs on ai coding agents from umass researchers. the model stayed the same. only the harness changed, the wrapper code around it. four different models tested. the finding: what you build around a model determines how well it actually performs in practice.
1
70
12-week build. autonomous ai research facility now runs real experiments across biology, chemistry, materials science.
1
50
space bunny's 1m context beats most frontier models. bigger than opus 200k or gpt-4 turbo 128k.
Not sure what model Space Bunny is, but I am impressed. I had a chance to test it this week. It's free in OpenCode for the next week, with a 1M token context window and image input. I focused my tests on visual work, mostly starting from a single image or sketch. Here are my results. Overall, its visual understanding is excellent, and its design taste is strong. One behavior that stood out is that it checks its own work. When it has a browser, it opens the page, looks at screenshots, and fixes what it sees before it says it's done. These verification capabilities matter in long-horizon agents. From a rough wireframe with seven handwritten notes, it built a full landing page and followed every note. From one photo of a moka pot, it built a 3D model that matches the original down to the eight-sided body, the brass valve, and the camping stove underneath. It built a 3D anatomy of a mirrorless camera with 67 labeled parts, an 11-element lens, a 9-blade iris, and an exploded view. Very excited about these results. I asked it for a 60-second animated proof of why a circle's area is πr². It built the whole thing, with slices that rearrange into a rectangle, a live slider, captions, and narration. Not perfect, but still impressive. I gave it a rough hand-drawn sketch of a game level. It built a playable 3D game in Three.js that follows the sketch and all ten rules I wrote on it. Then it playtested its own jumps in Chrome and tuned the physics until every lava stone was reachable. Based on this, I would reach for it for image-to-code work, interactive prototypes, and creative coding in 3D. It’s also quite fast, which makes it great for rapid prototyping. Space Bunny generated everything in the video in OpenCode. I wrote the prompts and supplied the images.
71
clm is the second system one model this month. runs 9x faster than jev while showing better verification at long-horizon tasks. matters if you're building agent harnesses that need to check work across extended reasoning chains. verification speed compounds on multi-step flows.
41
community already testing it on minecraft generation and video editing. both clearing expectations on first runs.
A mystery model just appeared on OpenRouter: Space Bunny Alpha. It’s free to try, has a 1M-token context window, accepts images and video, and is pitched as fast and strong at coding. Community fingerprinting points to a MiniMax M3.1 preview (@MiniMax_AI), but nobody has officially confirmed who made it. If that guess is right, this is a pretty fun way to preview what MiniMax has been cooking.
29
gpt-6 sol makes half the factual errors of gpt-5.6 while running at 50% lower cost per token.
1
53
opus 5.5: 74 tok/s at benchmark 58. gpt-6: 23 tok/s at 48.
1
1
60
ai made finding vulnerabilities free. it did nothing for the volunteers who have to fix them.
Release 2609 is live! In this Progress Report, we were harassed by AI (yes, even more), add an XFB resolution display for the curious, finally let you play Bird's-Eye Bull's-Eye on Android, added NetPlay to Android, and more! dolphin-emu.org/blog/2026/09…
123
google shipped gemini 3.8 tts yesterday. comes with 2,000 production voices covering 130 languages. you can design new voices from text prompts or clone a voice from 30 seconds of audio. scored 71.4 on hume ai voice design benchmark, taking first place.
1
1
60
950 agents searched 200,000 reverse transcriptases, 210 million tokens, 21 hours. system called ART, found in bacteriophages
First math. Now we’re starting to see similar acceleration in biology. it was fast... AI is starting to make big biological discoveries Anthropic says Claude autonomously discovered a previously uncharacterized enzyme system with CRISPR like DNA repeats while searching massive genomic databases. Around 950 Claude agents worked for 21 hours, analyzed more than 200,000 reverse transcriptases, generated 1000s of candidates, and surfaced a system human researchers then validated experimentally. The function of the new system, called ART, is still being investigated, but its architecture resembles other programmable biological systems capable of manipulating genetic information. If this kind of AI driven genome mining scales, it could compress weeks or months of expert research into hours. 👀 "This is the first result from our new molecular biology lab, where a team of Anthropic biologists is using Claude to explore and accelerate fundamental biology research. There, Claude works through data and literature to generate hypotheses and candidate biological systems to study. After our scientists review Claude’s hypotheses, they test the most promising ideas, with all lab work done by our scientists. We’d like to extend this approach to a broad range of problems—in genomics and in other fields" 1000s x acceleration incoming. because of that aging reversal therapies are near (2030s)
41
asked opus 5.5 to draw how a transformer reads a sentence. one prompt, under 10kb.
38
disclosed is doing all the work in that sentence. 14.09% is a floor, not a measurement.
Disclosed AI use in math rose from 1.39% to 14.09% in under 6 months. The authors scanned 32,944 math-related arXiv papers posted March 1 to August 20, 2026. Proof construction was the biggest category, appearing in 1,225 papers. AI is not spreading evenly across math: Combinatorics had the largest number of AI-assisted papers overall, while Metric Geometry had the biggest percentage of its own papers using AI.
43
people outside the labs already have a say. they switch models in an afternoon.
People outside the AI labs should have a real say in how this technology develops, and a clear way to judge if it's happening safely. Standards should help prevent the concentration of power, including by making sure new companies and open-model companies can compete. They should also help countries and companies compare evidence and learn from failures. We think the US should lead this effort. Here is our proposal: openai.com/index/building-st…
50
Jatin Garg retweeted
“Message market fit is when you have people signing up. Product market fit is when they retain.” Juicebox got 30,000 signups in 3 days. Then came six brutal months getting the product to actually work. David Paffenholz. Co-founder and CEO of Juicebox. Started the company at 21. 90 employees, 5K+ customers today. Bet on getting search right first while competitors tried to build the entire “AI recruiter.” Knuckle Up ↓ In this conversation with David: 0:00 Who is David Paffenholz? 2:42 Why betting everything on search beat the "AI recruiter" hype? 6:57 From search to copilot to agents: how has the product evolved? 10:02 Why hasn't LinkedIn already built this? 11:34 What's stopping OpenAI or Anthropic from just building this instead? 13:14 How does Juicebox stay ahead as competitors ship just as fast? 16:25 What made David's first 60-second LinkedIn video actually work? 18:12 "Message market fit" vs product market fit: what's the real difference? 19:12 What was the brutal six months after launch actually like? 22:50 Why are large enterprises still slow to adopt AI recruiting? 30:49 What are Juicebox's only two company values? 32:17 How does Juicebox mandate five days in office without going 996? 35:17 How has Juicebox never lost an employee in two years? 45:32 Does David think 90% of employees need to be "AI-pilled"? 56:57 What happens when a candidate gets 240 outreaches and answers four? 1:05:50 What was David's biggest mistake as a first-time CEO? 1:10:30 Quickfire: red flags, overrated advice, and Cursor's hiring machine 1:16:10 How is David thinking about kids and life at 25?
7
13
145
342,488
the numbers this week back that up. a browser agent finishing a task in 7 seconds for $0.0039. PR review 200x cheaper than claude, answered in half a second. computer use 155x cheaper than opus 5. none of it needed a bigger model.
What I like about Jev: for years, large generalist models sucked almost all the oxygen out of AI. But there are huge opportunities in much more specialized and customized models, built for specific tasks and languages and as a result, orders of magnitude cheaper, faster and more optimized. There are 3 million of them publicly available on @huggingface. Let's build a much more diverse AI ecosystem!
95
ctrl+F on the HN hiring thread has never worked. you search "remote" and get 400 posts that say "not remote". gave jev the same thread. 200 posts, 39 real matches, 2.5 seconds, half a cent.
75