builder, ai, markets, and general data nerd | dodatathings.dev

In the lab 🧪
Foliome now prices every goal in your household's plan and gives you everything you need to monitor and manage your wealth. Not just another dashbaord. Open source, runs on your machine.
Article

The Financial Plan Advisors Charge Thousands For, Free and Open Source

I just pushed two things to Foliome that matter to anyone managing their own wealth: the Personal Financial Statement and Periodic Statements. They're two halves of one idea. Most financial

1
2
441
Anthropic got all their users like
Anthropic is on a legendary run. Max and Team subscribers get $100/$200/$500 in API credit to test EVERY month anthropic.com/claude-haiku-5…
18
Small models ate office work. Haiku 5.5 edges Opus 5 on GDPval (1620 vs 1596) and AA-Briefcase (1578 vs 1562). Opus still wins HLE, 53% vs 45.9%. The flagship's remaining lead sits in hard-reasoning depth, and the work-task evals no longer pay for it. A team running document and spreadsheet workloads can default to Haiku at ~75% less cost and escalate to Opus only on the HLE-style tail.
Haiku 5.5 makes absolutely no sense lol. Anthropic somehow made Haiku MUCH smarter... then made it ~75% cheaper to run. and it's borderline better than Opus 5: Haiku 5.5 vs Opus 5: > GDPval: 1620 vs 1596 > AA-Briefcase: 1578 vs 1562 > HLE: 45.9% vs 53% > Terminal-Bench 4.0: 39.2% vs 46% yes... Anthropic's cheap, fast model literally beats its old flagship on BOTH knowledge-work benchmarks. what the heck.
2
2
73
Your laptop can become your next server. Agents that run continuously need a machine that stays on and keeps a 120B+ parameter model in memory. That's the pitch behind Surface Laptop Ultra, with the 1T-parameter DGX Station as the deskside tier of the same stack. Execution Containers are the sleeper in this list. An always-on agent gets an OS-level sandbox instead of your whole file system, so Windows now ships its own answer to what an agent can touch while you're away from the keyboard.
I will just leave this here 💚 Jensen said that software engineering has already changed - engineers are either building AI and improving models, or using AI to build things. That means the notion of what a personal computer provides needs to change. Builders will use local models, cloud agents, and hybrid, without even thinking about the complexity behind it. Some of the big announcements: 💚 Surface Laptop Ultra: A new NVIDIA-powered laptop capable of running 120B+ parameter AI models locally. 💚 NVIDIA RTX Spark PCs: A new generation of Windows laptops and compact desktops designed for personal AI agents. 💚 Surface RTX Spark Dev Box: A dedicated desktop machine for local AI development. 💚 NVIDIA DGX Station for Windows: A deskside supercomputer designed to run models as large as 1 trillion parameters. 💚 Microsoft Execution Containers: Secure, OS-level infrastructure for agents that operate continuously on your computer.
50
Same sticker price, 3x the tokens. Haiku 5.5 and GPT-6 Luna both list at $0.10/$0.50 per 1M. At max effort, Haiku burns ~162k output tokens per task against Luna's ~50k. That's roughly 8 cents of output versus 2.5 cents, for a 43 versus 38 score. At matched intelligence the gap closes. Haiku (high) scores 38 on ~55k tokens, so the extra 5 points cost about 3x. The 5x step to $0.50/$2.50 above 100k prompts hits long-context agent loops first. Artificial Analysis's provisional cost figures don't include it yet.
Anthropic has released Claude Haiku 5.5, scoring 43 on the Artificial Analysis Intelligence Index - up 26 points one year after the last Haiku release Haiku 5.5 is the first Haiku model with Anthropic’s effort settings and adaptive thinking, and Anthropic has introduced tiered pricing. Haiku 5.5 is cheaper than its predecessor - it costs $0.10/$0.50 per 1M input/output tokens for prompts up to 100k tokens (the same as GPT-6 Luna and 10% of the previous Haiku model). However, this pricing rises 5x to $0.50/$2.50 above 100k. The site does not yet reflect tiered pricing, so provisional cost figures for Haiku 5.5 do not include the step up cost. We are working on support and will follow up with Cost per Task coverage soon. Key takeaways: ➤ Leading small-class model performance: At max effort Haiku 5.5 sits slightly ahead of models such as GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38). Its score is comparable to Kimi K3 (44), a 2.8T parameter open weights model, and trails Claude Sonnet 5.5 (max, 56) by 13 points ➤ Heavy token use compared to GPT-6 Luna: Haiku 5.5 (max) uses ~162k output tokens per Intelligence Index task, ~3x GPT-6 Luna (max, ~50k). Moving from xhigh to max adds 2 points for ~1.8x the tokens. At similar intelligence it also uses more tokens than GPT-6 Luna: Haiku 5.5 (high) scores 38 with ~55k tokens per task against 38 with ~50k for Luna (max), and the gap widens at lower effort settings ➤ Highly capable at agentic knowledge work: on AA-Briefcase, our private evaluation for realistic knowledge work tasks, Haiku 5.5 (max) reaches 1578 Elo, ahead of models including Kimi K3 and GLM-5.3, and comparable to Muse Spark 1.3 (max) ➤ Improvements on terminal use: on Terminal-Bench 4.0 it scores 33%, up from 0% for Haiku 4.5. This is level with GLM-5.3 Flash, and ahead of Gemini 3.8 Flash (20%) and GPT-6 Luna (13%) ➤ Lower factual knowledge, but relatively low hallucinations: as expected for a smaller-class model, Haiku 5.5 has lower factual knowledge than its siblings. AA-Omniscience accuracy is 36%, against 55% for Gemini 3.8 Flash and 44% for GPT-6 Luna, but this is partly driven by more willingness to admit when it doesn’t know - its hallucination rate is lower, at 40% against 55% and 77% ➤ AutomationBench-AA result likely understated: Haiku 5.5 scores 35%, against 53–60% for GPT-6 Luna, Gemini 3.8 Flash and GLM-5.3 Flash. During pre-release testing, a safety refusal issue caused the model to over-refuse. Anthropic is working on resolving this - we will re-run this evaluation with the fix, and expect this score to rise Other model details: ➤ Context window: 1 million tokens, up from 200k for Claude 4.5 Haiku ➤ Pricing: $0.10/$0.50 per 1M input/output tokens up to 100k tokens, $0.50/$2.50 above. Cache reads $0.01 ($0.05 above 100k), 5 minute cache writes $0.125 ($0.625 above 100k) ➤ Multimodality: Text and image input, with text output
25
Search your phone, privately. One embedding model across code, images, video, audio and text means a single index can answer "the clip where she mentions the invoice" with no cloud call. At 270m to 740m parameters, it's sized for the device in your pocket. Matryoshka embeddings let you truncate each vector and keep it usable, so a whole camera roll's index can shrink to whatever storage the phone can spare. Apache 2 means app developers can bundle it without a license conversation.
Introducing EmbeddingGemma 2, our new open embeddings model for on-device use cases! 👀 Code, image, video, audio, and text embeddings 🤏 Modular, going from 270m to 740m parameters 🪆 Matryoshka embeddings for storage efficiency 🤗 Apache 2 License
31
Right outcome, unearned action. Across ten model-harness configurations, agents that judge well when asked whether an action is justified do much worse when they have to gather the evidence and execute it in a live workflow. Failures often start before execution. An agent can update the correct record on a guess and still get graded a success. Outcome-only evals count that run as a pass, so the pass rate overstates how safe it is to hand the agent write access. Teams shipping agents with write permissions should log whether the supporting evidence appeared in the trace before each write call fired. arxiv.org/abs/2610.07753
1
33
44% in a few months. Down 11% YTD in late March to up 28% now works out to roughly a 44% rally off the low. That's the $1.2T of added market value. $5.78T is only about 3.8% below $6T, so a normal post-earnings move clears it. Cap-weighted index funds mechanically hold more Nvidia after a run like this. Anyone in one now has a bigger slice of Nvidia's next earnings print than in March, whether they chose it or not.
Nvidia is on the verge of making market history. $NVDA is now valued at $5.78 trillion, putting the first ever $6 trillion public company within touching distance. What makes this run even crazier is that Nvidia was down 11% for the year in late March. Now it’s up 28% in 2026, adding around $1.2 trillion in market value along the way. From negative YTD to knocking on the door of $6T in a matter of months. The AI trade is still showing no signs of slowing down.
20
This is my AI Engineer going on call
Ben Affleck reveals he writes Python, understands convolutional neural networks, worked extensively with GPUs, and used his celebrity status to get private looks at Google and OpenAI’s video models “I’ve always been kind of into computers since I was young. Then, when film started to move from analog film to digital, I became more interested in that aspect of it. The visual-effects workflow for many years has included machine learning, so I can write pretty shitty Python scripts and stuff like that. “With convolutional neural networks, which were the precursors to what the transformer can do, which is much more computation simultaneously, you would do things like look at what’s called a tensor. That’s the numerical translation of a visual image in numbers, like the batch number, the frame number and the red, green and blue values of each pixel in each frame. It’s just that simple. That numeric is called a tensor. “You’d use a convolutional neural network to identify patterns that reveal what’s called edge detection or feature extraction, which is identifying patterns well enough to know, this is where the window ledge is, so we can more easily take the green-screen image out and replace it with something. “That was familiar to me early on because, prior to Artists Equity, I had a small visual-effects company. I’ve worked with GPUs a lot too. The visual-effects guys said, ‘Hey, you should see. There are a couple: Google and this other company, OpenAI, are doing really interesting stuff with transformers in video.’ “I’ve learned that I can actually just call up and go, ‘Hey, it’s Ben Affleck. Can I come see what you’re doing?’ Sometimes people say yes, to my astonishment.”
43
As fast as decision models just got here: they just halved in price. At $0.02 per million input tokens, a 1,000-token call costs $0.00002, so a dollar buys 50,000 decisions. v1 sat at roughly double that. That price makes a model call on every event cheaper than building a sampling scheme to skip most of them. The weights are open at 27B, so any inference host can serve the same model and bid under Perplexity's rate.
pplx-decider-v1.1-27b, our updated open weights multimodal decision model, is now available. It scores the highest on the new @huggingface Decision Index 0.3 benchmark. pplx-decider-v1.1-27b costs half as much as v1, at $0.02 per million input tokens. huggingface.co/spaces/multim…
1
1
54
Metrics become the engineering job. Wang's recipe has two parts: an agentic loop and an evaluation metric the agents optimize against. The code is the cheap part. The loop is only as good as its scoring function. Work with a clean automatic score, like a test suite, a latency target, or a build pass rate, falls to swarms first. Work where "better" is a judgment call stays with humans longer.
Alexandr Wang, Meta's chief AI officer: "Internally at Meta, we have seen cases where, if you develop the right agentic loop and have the right evaluation system and metric for the agents to optimize, a swarm of agents can accomplish more than a team of 100 engineers. They can do it very handily, actually, very easily." Wang was in conversation with Garry Tan, head of Y Combinator, at Startup School 2026. Interested in AI? Follow @AiEvolutio58513 and never fall behind. I track ChatGPT, Claude, and every tool quietly changing how we work and create, then hand you the tested signals.
1
32
Text watermarks just went mainstream. OpenAI will mark eligible ChatGPT and Codex text in the EU over the coming weeks. Text watermarking typically nudges token choices toward a statistical pattern that a detector can test for later, so nothing visible changes for the user. Run the output through another model to paraphrase it and the pattern mostly washes out. Codex is on the list too, and code leaves the fewest free tokens to nudge, since syntax forces most of them.
OpenAI announces that eligible text generated by ChatGPT and Codex will begin being watermarked in the EU over the coming weeks to comply with the EU AI Act.
43
This open-source repo automates one of the most complex pieces in all of wealth management: the personal financial statement. The PFS sits at the center of wealth planning. It's where everything meets: account balances, paychecks, taxes, the mortgage, college in 9 years, retirement in 24. Each piece comes from a different place and rests on its own assumption. The PFS weighs what the household will earn against what it has promised to spend, and puts a price on each goal. The fundamentals are the same for every household, and it's beautiful when put together. But it is a nightmare from a data monitoring perspective. Here you get it for free.
1
75
Agents miss the same bugs. Claude Code and Codex reviewing each other's patches lifted fully correct fixes from 45.8% to 62.5% at matched spend. A bigger budget for one agent mostly buys more passes through the same blind spots. Silent bugs are the ones the tests don't flag, so a patch can pass CI and still be wrong. Models from different training lineages fail in different places, which makes their disagreement the useful signal. A stack where one vendor writes the patch and the same vendor's bot reviews it is paying twice for one perspective.
New Meta paper finds that having 2 different coding agents review each other's patches catches far more silent bugs than giving 1 agent a bigger budget. Mixing Claude Code and Codex on the same task raised fully correct patches from 45.8% to 62.5% at matched spend. Agents editing real training code can leak test data, break a gradient, or miswire a flag. The code still runs, so you burn GPU hours and get numbers that look valid. The reason is uncorrelated mistakes: agents from the same product fail the same way, so there is little for review to catch. The effect repeated on an unrelated training codebase. – arxiv. org/abs/2609.39551 Title: "RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models"
32
Checkout pages are for humans. Stripe moved $1.9T in 2025, up 34%, on rails built around a person typing a card number into a form. An agent buyer breaks that. Fraud models are trained to flag bots as attackers, and a CAPTCHA is a wall to a shopping agent. The new problem is telling an agent acting for a real customer from a scraper with a stolen card. Stripe's fraud systems have that $1.9T of labeled payment history to learn it from.
NEW: Stripe Co-Founder John Collison (@collision) says AI agents will completely rewire internet commerce. @stripe processed $1.9T in 2025, up 34%. Now the company is rebuilding for an internet where AI agents, not humans, increasingly transact. “Instinct is real, Muse is real. All these things are real.” “Computer use is the ultimate backwards compatibility layer for the real world.” I sat down with John at @browserbase Navigate to talk agentic commerce, computer use, cybersecurity, how Stripe is rebuilding for AI & why Stripe considers 2026 the beginning of the “singularity epoch.” John sees AI already showing up in new company creation. Stripe is seeing its highest rate of new firm creation in recent memory, while its 2026 Atlas cohort is tracking at 5x the revenue of the 2025 class. We also cover Stripe’s recent acquisitions: › Bridge + Privy, bringing crypto-native infrastructure & talent › Metronome, usage-based billing for the AI era › OpenRouter, routing across multiple AI models as inference spend & model competition explodes Browserbase is building the infrastructure for this shift, giving AI agents cloud browsers to navigate the existing web & take actions across it. 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Stripe’s $1.9T & Rebuilding for AI (00:31) Why Agents Will Replace Web Forms (02:20) How AI Rewires Commerce & Product Discovery (04:40) More Cyberattacks? (06:49) Why Stripe Says the Singularity Started in 2026 (09:29) Stripe Is Rebuilding Its APIs for AI Agents (11:38) Why Stripe Bought Bridge, Privy, Metronome & OpenRouter (15:12) When Stripe Builds vs. Buys (16:55) How AI Is Changing Stripe Internally (20:39) How John Keeps Up With AI (23:28) The Coming 'Good Bot' vs. 'Bad Bot' Problem (24:53) 17 Years of Lessons Building Stripe
57
This is more amazing when you think how Anthropic projects profitability vs. OpenAI. Makes you wonder just how @AnthropicAI does it. They do have a good product.
Holy sh*t, Anthropic’s subscriptions offer so much more value, even compared to the new and efficient GPT-6.1 Sol. It’s not even close. SemiAnalysis puts Opus 5.5 at over 5x the API-equivalent value of GPT-6.1 Sol for agentic workloads at the same subscription price. It’s not even close on this measure. And on top: OpenAI just halved the usage limits on its $200 ChatGPT plan.
1
70
GPUs age faster than track. If AI capex really is bigger than the railway boom, the comparison cuts both ways. Railway investors got wiped out in the bust, but the track kept hauling freight for generations, so the economy kept the asset. AI capex buys chips and servers that go obsolete in years. An overbuild this size would not leave a durable asset behind. It would leave write-downs.
The ongoing AI capex boom is shaping up to be the largest on record, bigger even than the railway boom. Source: brookings.edu/wp-content/upl…
1
51
Running a fab is the hard part. Pouring concrete in Texas and buying tools is the purchasable piece. Yield, the share of working chips per wafer, comes from years of process tuning, and TSMC holds that know-how. Talking to TSMC about operations means Musk is shopping for it. An operator usually brings its own process recipe. A TSMC-run Terafab would likely produce on TSMC's process, so Intel's 14A claim shrinks to whatever volume Terafab sends it separately. Musk's own description is "just discussions.”
$TSM is up and $INTC down premarket after a report TSMC is exploring helping run Musk’s Terafab fabs in Texas with Musk saying talks are “just discussions but something may come of it.” This does not necessarily push Intel out since Terafab can still be a major 14A customer but it gives Musk more flexibility for $SPCX and $TSLA across both foundries as chip demand scales. This would be a major win for TSMC if Musk provides the anchor demand for a bigger Texas buildout while TSMC monetizes its manufacturing expertise without funding the entire fab itself.
1
45
LeCun tells academia to skip LLMs. His line at ETH Zürich: "There is nothing you can bring to the table." Frontier LLM work now runs on pretraining budgets a university lab can't fund, so a grad student's marginal contribution there is close to zero.
Yann LeCun's latest talk at ETH Zürich "You should not work on LLM. At least if you're in academia, you should absolutely not work on LLMs. There is nothing you can bring to the table. If you're interested in making real progress in AI, in sort of grounded AI for the real world, if you want physical AI, don't work on LLMs, and don't work on generative models either. So, as you can probably guess, this does not make me very popular in Silicon Valley." ---- From "Perfology Clips" YouTube channel, (link in comment)
1
2
183
Six months behind the frontier. Underdog's Law is a testable claim: what frontier models do today runs on a phone or laptop you already own within six months. The mechanism is co-designing the model and the inference engine together, so the hardware gets used fully. If it holds, a capable assistant stops being a subscription and a data handoff. Your email, health records, and calendar never leave the device. Free and private also rules out the usual ad-funded payback. That puts the revenue answer in the agentic commerce and internet economy research they list.
Introducing Underdog, your Private Personal AI on devices you already own Today we're announcing our backing from @a16z @khoslaventures @HummingbirdVC Anthology (@AnthropicAI @MenloVentures) @patrickc @naval @rauchg @polynoamial @Thom_Wolf @Mascobot @tszzl @OfficialLoganK and other top AI leaders Our mission is to provide free, capable, reliable and private AI to billions of people. Underdog’s Law: today’s frontier intelligence reaches your devices in six months. We’re starting with fast, capable AI that runs on consumer hardware. By co-designing models and inference engines, we’re pushing the frontier of capability, speed, power and data efficiency. Our research also spans agentic commerce, confidential inference, and how AI will reshape the internet economy. Privacy and capability no longer needs to be a tradeoff. If you believe in this future, join us. Time to build.
2
2
117