Co-founder, CTO, designer @goodygifting. Technology and behavioral science for social good. Systems, mindfulness, psychology, writing, and living well 🪴

San Francisco
Opus 5.5 is my favorite model yet So many engineers on my team saying how much they like it I've been using it for projects all day and it's incredibly capable, basically one-shots most things, and it sips usage... incredibly efficient Monumental vibe shift from Fable anxiety
1
5
304
the pace of stuff is whiplash inducing. not long ago was plan mode the gold standard best practice for serious agentic coding
we’re thinking of killing plan mode and using the shift+tab hotkey to adjust effort levels I don’t think the models need plan mode anymore, but if you’re a plan mode diehard would love to get your feedback on why
2
4
683
and no I don’t use it
1
76
a tweet like this describing this in such detail, even jokingly, would only be posted in a low trust society
Hello computer, Find every Amazon purchase I ever made. Call the number of the manufacturer. Say you’re not satisfied. Request a refund. If they decline, threaten to leave a bad review. If they decline, find the CEO’s phone number and ask him for a refund. Now, repeat the same process for items I never even purchased and see if you can get them to send me money. If they ask for a recipient, generate an image of one and send that. Do not stop until you’ve collected $1 million in refunds.
1
3
923
these costs have gotten truly insane
My project manager came into my office this morning to tell me she’s stepping down to become a stay-at-home mom. She makes $68,000 a year. Her husband earns $72,000. They just had their second child, so she called licensed local daycare centers to get quotes for an infant and a toddler. Average cost: $2,650 a month. That’s $31,800 every year straight out of their post-tax income. After federal taxes, state taxes, and healthcare deductions, her take-home pay was roughly $3,700 a month. After paying $2,650 for daycare, she was working 40 hours a week to net $1,050 a month…..or about $6.50 an hour. She looked at me and said: "I’m literally spending 160 hours a month away from my kids just to clear $250 a week after childcare costs." We didn't just build an expensive economy. We built a system where holding down a mid-class job as a second earner costs more than staying home.
398
this is painfully good the recursive self upgrade scene is especially masterful
Claude Opus 5.5 has the best visual design of any model I have tested so far
325
We just opened our first showroom in San Francisco! 🎉 Come visit Goody at 330 Pine Street (near USQ and FiDi), and we'll also be hosting events through the rest of the year!
4
2
31
2,849
It's such a dream come true to have a retail space, it's a bit surreal 🥹
4
131
and that's why tobi is the goat
We are excited to announce we are partnering deeply with Muse to enable agentic checkout with Shop Pay on all Shopify stores, offering people an easy and delightful way to shop and check out with Muse.
1
1
376
Mark Bao retweeted
i think one of the most romantic things about being alive is that another person can alter the architecture of your life so completely that new versions of you begin to exist. not every part of you is waiting to be uncovered. some are formed through being known by someone else
47
5,492
28,882
461,323
A lot of people are missing Terence Tao’s point and thinking “mathematicians are upset that AI is better than them.” That’s not what he’s saying, and some people are forgetting that Tao is one of the most AI-pilled mathematicians out there. His point is that when people work on discovering something, along the way they invent new concepts. Those concepts later become useful far beyond the original goal, and enables further inventions. Finding a solution does matter, but the intermediate idea is often what makes the field richer, because other people can share it and build the next thing from it. In tech, we can use the analogy of collaborative software. We started with algorithms for merging changes in a Word document, and evolved that to concepts about versions, diffs, and merges, and later to real-time collaboration tools like Git, Google Docs, and Figma. Humans built upon these concepts and developed more powerful solutions. Terence’s worry is that a machine automating a solution robs the field of the value of developing the intermediate discoveries in the pursuit of larger discoveries. When automating a solution, the intermediate discoveries and invention of concepts can be buried or completely hidden in the black box. We don’t learn from them to build the next thing; it’s like we never made the invention of collaborative document editing and thus could not have the conceptual understanding to invent the next version – and since it’s hidden, we also don’t socialize them to allow other people to invent, too, a core tenet of collective discovery. So then, in both code and math, this leads to the atrophy of development of concepts in the field. In other words: pure ‘solution extraction’ that hides the process of discovery can leave the field with a checked-off theorem but little new insight or new questions to pursue. And it might prevent us from understanding a field deeper. I am seeing, first-hand, that atrophying of skills in software development. We push buttons and get solutions. There is much less incentive to develop new concepts and human skill. The bet most software companies are making is that LLMs are so effective in writing code that you’re still shipping overwhelmingly more value even with human skill atrophy, and it’s the right bet IMO. However, much of the software industry is built upon building things, not necessarily novel invention and research. In such an environment, you can say that you accept some atrophying of conceptual invention and human skill for more output. On the other hand, sectors like math and pure sciences that are focused on invention and insight might be the hardest hit by this. Practical/applied sciences might fall somewhere in the middle. An Alzheimer’s cure, room-temperature semiconductor, or highly effective carbon capture solution are far too valuable to sandbag and say only humans can do that to develop concepts in the ‘proper’ way. The outcome matters too much to treat the preservation of concept invention as the highest goal. Even there, though, hidden intermediates can slow the next breakthrough if nobody can see how the first one actually worked. So the question is not “is AI allowed to solve hard problems?” It is “in this field (math, science, tech, etc.), is the answer itself the main point, or are the concepts and abstractions we use to get there also the thing we need to maintain?”​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​ In pure math, there’s an argument that the intermediates are often more useful than the solution, and atrophy in concept development is highly detrimental to the field. Solving Navier–Stokes, contrary to what some people claim, has little practical application, and pure math might be one of those fields where just finding a solution isn’t the entire point, and can actually be contrary to the field, which is what Tao is worried about.
as predicted, Terence Tao and friends are not happy Mathematicians that spend their lives trying to solve math problems are unhappy that AI is solving math problems because how dare they!
333
723
3,744
470,359
Whether you agree that research/pure math should be 'special' in this way is arguable. In my opinion, research math sometimes becomes applied math and affects the real world, and given that I think it's worth advancing the world at the sake of research math. But it's not as simple as "mathematicians mad" and we would be better off engaging with the actual problem.
9
4
67
8,375
ok they're cooking
Introducing Projects, a new way of working in Cursor. Rather than creating a chat for every task, you work with a coordinator agent in a single, persistent thread. Like @bot, your agent is always on, proactively manages work with subagents, and improves over time.
1
857
screw Millennium Problems and antibiotics or whatever. imagine 10,000 agents discovering variations of ‘you’re telling me a shrimp fried this rice?’
astra benchmark performance: still pretty poor otoh; “self raising flour? where are the parents?” did get a modest chuckle
1
2
599
Mark Bao retweeted
GPT watching me ask which air fryer to buy while OpenAI’s other agents are solving Millennium Prize problems
54
1,092
20,488
355,659
Mark Bao retweeted
As recently as April this year, prediction markets gave AI less than a 40% chance of solving any Millennium Prize Problem before **2030**
36
111
1,318
73,648
Navier–Stokes is cool but it has limited practical applications in real life A room temp semiconductor though.... now they are cooking
Replying to @anabology
you know what let's try
1
4
495
please let this shitpost be the reason openai spins up a cluster and discovers a life-saving medicine, it would be so fkn funny
Rumour is that anthropic has discovered a totally new class of antibiotic effective against resistant bacteria. They plan to announce it at their IPO. They really hope a competitor doesn't spend a bunch of compute to try and scoop them and release it first......
6
752
And still, today we are standing near the lowest point of the exponential curve, compared to what the future holds...
In 3.5 years, we went from GPT-4 which couldn't reliably add two numbers together to a 10,000 agent system that can autonomously solve a Millennium Prize problem. This is what exponential progress looks like. Accelerating, awe inspiring and terrifying.
2
392
insane that we are still seeing jumps like this
Replying to @OpenAI
This model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations.
3
372