We find the money hiding in your data Amazon, Crypto.com, Morpho Delta Labs · deltalabs.fun

about 1 in 5 a/b tests that come back significant at 95% do nothing when you ship them. the stats aren't wrong, the problem is what you're testing: around 70% of it moves nothing. and a 5% error rate on a pile that large produces plenty of fake winners: ron-berman.com/papers/fdr.pd…
1
12
I have a pretty simple suggestion for Dario: If you want to slow things down, just move Anthropic to the EU.
248
433
6,214
190,947
new paper: a zero-shot foundation model beat every bespoke electricity price forecaster on accuracy, then lost to the bespoke one on actual battery arbitrage profit once risk mattered. better forecast, worse decision. the loss function clearly is the model here. arxiv.org/abs/2609.00089
2
35
kapoor and narayanan surveyed ml results across 17 scientific fields and found data leakage in every one, 294 papers total. in their own replication of civil war prediction, the complex models stopped beating plain logistic regression as soon as the leakage was fixed: cell.com/patterns/fulltext/S…
1
26
Be honest, who panicked a little today?
2
36
DeltaLabs retweeted
All the AI is down
570
269
7,421
1,076,419
a new paper shows that 28-69% of accounts flagged as "at risk" by typical churn models were never actually churning. the model was just confusing seasonality with decline: deltalabs.fun/blog/seasonal-…
2
44
👀 Scaled enterprise AI returns about 7% (risk-free T-bills currently pay 4.1%), under the ~10% cost of capital most firms use as a hurdle. Top decile clears 18%. The largest gap between them isn't modelling: 68% of the leaders have mature data and governance practices, against 32% of everyone else. nobody's ROI problem was the model. ibm.com/thought-leadership/i…
1
34
📢 4 GitHub repos worth cloning this week if you build with Claude or LLM agents: 🔹 𝐚𝐰𝐞𝐬𝐨𝐦𝐞-𝐥𝐥𝐦-𝐚𝐩𝐩𝐬 (118k stars): a cookbook of 100+ ready-to-run AI agent and RAG templates. Every one is hand-built and tested end-to-end, covers the full modern stack (agents, RAG, MCP, voice, memory), and works across Claude, Gemini, GPT, and Llama. Apache-2.0, so you can fork it and ship it. 🔹 𝐏𝐢𝐱𝐞𝐥𝐑𝐀𝐆 (7k stars): a visual RAG framework that skips text parsing entirely. It renders documents, whether web pages, PDFs, or images, as screenshots, embeds them with a vision-language model, and serves search over a FAISS index. Benchmarked against Wikipedia's 8.28 million articles, and it ships a Claude Code plugin that lets Claude take a screenshot of any page and read it visually. 🔹 𝐛𝐨𝐨𝐤-𝐭𝐨-𝐬𝐤𝐢𝐥𝐥 (8.8k stars): turns any technical book PDF into a Claude Code skill. Run one command and it builds a chapter index, glossary, and cheat sheet, then loads chapters on demand when you ask about a topic instead of dumping the whole book into context. 🔹 𝐢-𝐡𝐚𝐯𝐞-𝐚𝐝𝐡𝐝 (4.5k stars): a Claude Code skill that stops the model from burying the answer under three paragraphs of preamble. Ten rules, one markdown file: lead with the action, number the steps, restate progress, skip the "hope this helps." - All four are free, open source, and install in under a minute - Worth trying if your Claude Code setup feels slower than it should #ClaudeCode #OpenSourceAI #LLMEngineering #RAG #AIAgents #DeveloperTools
3
17
89
9,325