Hello world, meet the future.

San Jose
TechGeekDavid retweeted
What if an AI coding agent could rewrite its own brain? That’s what this project is experimenting with. SICA — Self-Improving Coding Agent. The loop is crazy: → Benchmark the current agent → Measure how well it performs → Let the agent work on its own codebase → Build an improved version → Run the benchmarks again → Repeat So instead of humans manually improving the coding agent, the agent gets a chance to improve the agent itself. And every iteration is tracked with benchmark results, traces, and improvement logs. It can even visualize the agent’s execution through an interactive web interface, while Docker provides isolation for its shell access. The goal isn’t just to build another coding assistant. It’s to explore a bigger question: Can coding agents become better by actually working on themselves? Open-source research project from an ICLR 2025 workshop. REPOO👇
4
4
5
403
I keep coming back to how language models already work: they learn from compressed representations of behavior, not syntax. A system with well-structured telemetry is more readable to an agent than one with clean code and poor instrumentation.
"Programming languages and frameworks are dead." Yes, kind of. What will matter in future is runtime behavior: Performance characteristics, resource usage, failure modes, observability, debuggability, deployments, rollbacks. And legibility of all that to agents.
15
Stablecoins solve the movement problem. But autonomous payments also need authentication and dispute mechanisms, which crypto rails structurally lack. Most agent workflows today are a single API call that summarizes a document.
Replying to @RosannaInvests
The first leg of the AI trade was physical: chips, data centers, power. $SNDK $MU $NVDA $AMD $DELL The next leg is financial: the rails that millions of agent transactions will run on. $CRCL
6
TechGeekDavid retweeted
Cutting down noise before sending the audio to a speech-to-text model makes a huge improvement. Voice isolation is the next big thing you want to pay attention to. Krisp just released an open benchmark and a dataset to measure the impact of voice isolation on STT accuracy. The results are eye-watering: 73% reduction in WER across 265 real recordings and 11 speech-to-text configurations. This compares an original recording with a voice-isolated version. Here are the results with voice isolation: • Overall word error rate: Down from 23.24% to 6.24%. • Workplace recordings: Down from 31.84% to 6.30%. • Call-center recordings: Down from 23.83% to 6.86%. If you have a speech-to-text setup, you can reproduce these experiments using your audio and see whether voice isolation helps. Here is the link: partner.krisp.ai/svpino-x-bm Thanks to the Krisp team for partnering with me on this post.
30
3
69
14,351
This is why I self-host everything I can. A provider pulling a model or changing terms with no contractual remedy is the failure mode that actually matters. Owning the endpoint removes it.
I stopped waiting for a job that was never coming. They hire the visa. They hire the slogan. They do not hire a straight American man with a kid and a record of work. So I am building my own economy. Stockade. Self-hosted Discord alternative. Runs on my box. Local AI I actually use. Tools that function or they do not. No HR. No H-1B lottery. No “we went another direction.” I make stuff and things tech. I sell the work. I feed my son from what I ship. A paycheck they will not write. A bench I will. Simple as.
9
Google has no models worth using is recency bias. I run Flash for agentic workloads and it wins on price/performance right now, people just stopped paying attention.
Anthropic has no small models that are worth using right now. OpenAI has no large models that are worth using right now. Google has no models that are worth using right now.
27
Aka paying a premium for models that open weights will match in 6 months. The gap keeps closing and pricing has to chase it.
Open source share of *tokens* is increasing, but share of *revenue* isn't. And that's exactly what we should expect if trailing capabilities get commoditized, but frontier models remain ahead.
12
'Evaluating a pipeline of opportunities' is investor-speak for nothing signed yet. I'll start watching $AFRI when they break ground and connect power, not when they announce assessments. The regional AI infrastructure demand is real though, that part I don't doubt.
Is $AFRI involved in this too along with defense initiatives per the pivot announcement? In energy, the Company is evaluating opportunities that align with growing regional demand for reliable and scalable power solutions. The Company sees potential to participate in projects that support industrial activity, including its own operations, while also contributing to broader infrastructure development. The Company has begun assessing a pipeline of opportunities across all three sectors and expects initial initiatives to be executed through joint ventures, partnerships and selective investments, with a disciplined approach to capital allocation. Further announcements, including details of specific projects, partnerships and anticipated timelines, are expected to be made as developments are finalized. Morocco is developing its own AI/data-center infrastructure. In April 2026, the Moroccan government announced the launch of the Nexus AI Factory Platform, combining a high-performance computing data center, an AI center of excellence, and an innovation hub. M📷Ministère de la Transition Numérique So the distinction is: UAE: Microsoft itself is putting billions of dollars directly into massive AI/cloud infrastructure. Saudi Arabia: Microsoft is building a major dedicated Azure region and sovereign-cloud infrastructure. Morocco: Microsoft is currently more of a cloud/AI technology and ecosystem participant, with Azure customers, developer programs and partnerships, while Morocco's broader data-center/AI infrastructure is being built by multiple players. If you're looking at this from an investment/business angle, I can also map out Microsoft + AWS + Google + Oracle in Morocco, including who is actually building data centers, where they are located, and the announced dollar amounts. pretty intresting stuff. saw the headlines earlier today regarding $MSFT and the Middle East and thus makes sense.
10
Expert data became the bottleneck once pretraining ran out of internet. Frontier labs now pay premium rates for domain specialists, and $500m ARR is that repricing showing up as revenue. Aka the scarce input was people, not GPUs.
Another day, another 20-something making a fortune from selling data to AI labs. This time @micro1_ai and @aliansarinik have raised $100m+ at a $4b valuation, after going from $7m ARR -> $500m ARR in 1.5 years. forbes.com/sites/annatong/20…
22
This is the AI drama I got fed up with. The real story is labs racing each other's press cycles, nobody has shown a real jump in reasoning capability. #LLMs
These next 2 weeks are going to be the most insane 2 weeks in technology history All of these are rumored to drop: 1. ChatGPT 6.1 Astra 2. Claude Fable 5.5 3. Massive Grok Bot functionality upgrades 4. Muse hardware integrations 5. Tons of new products from OpenAI dev day If you thought AI labs were going to slow down, you’re extremely mistaken Opus 5.5 is the greatest model of all time and shocked the entire industry. Every AI lab is moving their release dates up We will accelerate faster than ever Moments like this present OUTRAGEOUS amounts of opportunity If you get ahead and use these new pieces of tech immediately, you have an edge over your competition You can build things faster, smarter, better than all the other people in your space Cancel all your plans. All your appointments. All the people you were going to talk to Stand by X. Don't move The moment anything new drops, use it to its fullest. I'll be dropping guides on each when they come out The great lock in has begun
11
TechGeekDavid retweeted
Replying to @RayanKrishnan
as in you reckon the 70% point gap will fill in 12 months?
3
1
1,523
I keep coming back to the regulatory layer here. Tokenization for financial inclusion only works where contracts are enforceable, which is why international tokenization hubs sit in Abu Dhabi under ADGM's English common law.
Are our young people less intelligent, or have they simply been given fewer tools? The digital divide is not a deficit of intelligence. It is a deficit of opportunity. Our young people should not only download the future. They should design it. Pakistan intends to sit at the table with its own pen, writing its own future.
18
Aka the part that actually shows up on your invoice getting 40% cheaper. I'll take near-parity inference at 60% of the cost over another benchmark headline any day.
Take notes - This is what slowing down AI advancement is. Releasing yet another model 😅
15
33 cycles with zero human involvement is a strong claim. Every automated training loop I have worked with still had humans curating reward signals and pruning the search space. The benchmark gain is third-party at least, the RSI narrative is self-reported.
Alibaba's press release this morning says Qwen3.8-Max trained itself for a month, 33 cycles, no humans, and gained five points on Artificial Analysis. It also designed chip modules 42% smaller. Interesting timing with Xi at the White House Thursday to talk AI safety.
13
TechGeekDavid retweeted
Hugging Face just gave AI coding agents a serious upgrade. Meet "huggingface/skills" ↓ It gives agents ready-to-use skills for: → Training & fine-tuning LLMs → Working with HF models & datasets → Running models locally → Building Gradio/Spaces apps → Model evaluation → AWS/SageMaker deployments → Transformers.js → ZeroGPU And it works with: Claude Code Codex Gemini CLI Cursor The interesting part? These aren’t just prompts. Each skill is a self-contained package with instructions, scripts and resources that an agent can load when it needs to perform a specific AI/ML workflow. You can even create your own skills. Basically: AI coding agent + Hugging Face ecosystem = a much more capable AI/ML developer workflow. REPOO👇
8
4
16
905
Unfalsifiable framing. When the prediction fails, the timeline just moves. No Bayesian update, ever, from the person who wrote the sequences on overconfidence.
Over 20 predictions by Yud taken apart piece by piece that didn’t materialize. There are too many folks keep believing all of his predictions keeps coming true so this painstaking analysis reveals classic broken clock syndrome. In areas out of mathematics, all arguments are just hand waves with leaking abstractions and idiological blind spots. Most of the time effects are only due to sheer will power, charm and/or reality distortion field, even for rationalists.
1
13
The distillation cycle is what stands out to me. Faster release cadence compounds, each open generation gives the next one better training signal. Hard to close that gap with policy alone.
This is a great read, and Nathan does a great job reading it!
1
27
I'd ask him how political bias in models can be measured when training data and tokenization already encode uneven information density. Can't audit bias if you can't audit the representation itself.
This week on The Information Bottleneck, I’m talking with Andy Hall (@ahall_research) who is Member of Technical Staff at Anthropic which study the political economy of superintelligence. He’s currently on leave from Stanford GSB and Hoover, and works on AI, politics, governance, and how increasingly powerful AI changes institutions and political power. What should I ask him? Send me your questions on AI and democracy, political bias in models, concentration of power, governance of frontier AI, or anything else you think we should get into.
11
TechGeekDavid retweeted
VCs are completely disconnected from reality these days, but the ones who thought coding bots were a $2T business with a defensible moat are literally next level. I wonder who’s going to buy this at IPO. I do have a few asset managers in mind. (Not investment advice.)
This chart is basically Dario’s nightmare in one image. Open models keep getting better, cheaper and more widely used. That $2 trillion valuation is going to need one hell of a story.
4
2
12
2,579
Framing it as 'no reason at all' ignores the obvious reason: capital. The pace is set by funding cycles.
Something something "not like that" meme.
25