building daydream: chatgpt's computer history, but for claude (OS) Math @UChicago

New York City
ChatGPT got Computer History. I built it for Claude and Cursor. DayDream remembers your apps, docs and pages on Mac, so you can ask “where did I leave off?” Free. Open source. Activity stored on your Mac. getdaydream.app
1
2
10,543
57
Your competitor is a VC-funded team of 18. You're 18 and too busy shipping. So we built DayDream: it understands what you're doing, so your agents know you better. Try it free 👇
1
1
2
45
Nicholas Fong retweeted
Your competitor is a VC-funded team of 18. You're 18 and too busy shipping. So we built DayDream: it understands what you're doing, so your agents know you better. Try it free 👇
1
1
2
45
Hypothesis: Right now Claude 5.5 is a beast and is destroying OAI. They want to regain supremacy so they are freeing up training compute my making their subs worse and increasing revenue by making API more attractive. By freeing up so much subscription compute and with the claude sub such a good deal, it will push a lot of people to claude subs. Claude is already compute constrained and this influx will force them to reduce compute used for the next training run, pausing new subs, or reducing API spend (which is really bad). Anthropic will thus be forced to lower usage for the sub as well. Which is what they have been intending on doing for a long time now.
2
131
It is of my opinion that Fable 5 is a significantly larger model than GPT Astra (~1.5x total params 2-2.5x active). That being said the pertaining cost of Fable 5 was minor due to generalization because of the size of the model while the pertaining cost of Astra was MASSIVE. This means that both have high costs and leads to each having different incentives for serving. Astra needs to be served at massive scale while Anthropic can only serve a moderate amount of capacity. That said this is entirely speculative just based on how each company is acting. My analysis is neglecting a key point: the cost of Open AI compute is much cheaper than Anthropic because they bought it up earlier. This also factors into how the models are served and the business strategy for each company. Cheers to both companies though on their remarkable success and their efforts in pushing the frontier of intelligence and safety. The medical and societal implications of this are so positive.
1
7
989
UMMMMM.... I don't know how I feel about my ASTRA agents being so good at communicating. They are able to multi-agent communication REALLY well. Never seen anything like that from any model previously. Enabling crazy amounts of work to be done, but also... yea.
4
237
Genuinely I don't know how it is possible that GDM has such misaligned RL. It is truly incredible. Clearly they just realized that the reward gradient could be negative for non-truthful answers. Maybe they should discover that it can also be negative for any code changes that do not lead to a better final result. As in introducing new bugs needs to be punished. Really don't get why they can't get their act together. I know RL environments are hard, but you have so many talented engineers that know what good code looks like. What you need to do is look at the cases. Then look at what we don't like. Try to generalize that into a reward function and then roll that out across all your environments. Like DUDE I am so frustrated with the misaligned behavior of gemini models. Now that we actually use AI for non-slop stuff we need to punish non-true/verifiable behavior later in the reward. It doesn't matter if it scores higher by increasing perplexity through your SHITTY RL if you can't trust it to make simple decisions. I am not saying to NUKE the perplexity but later stage reliability RL is NECESSARY. You are supposed to do it in alternating stages. With it essentially being a growing cone until you reach deployment. I am soooo frustrated because I love you GDM and I <3 your mission and I love your team. You do so much incredible work for the community and its a shame that your models are not up to par right now.
2
321
Holy SHIT I am so so so impressed with GPT Astra. Genuinely feels like another opus 4.5 moment!!! Step change difference in what you can do. This time (unlike mythos) we get limits too!
1
98
Nicholas Fong retweeted
AI discourse has been better than usual this week. I think the most annoying posters are all at Burning Man
23
13
687
37,091
Super excited with GPT Astra. Believe that MAINLINE capabilities (ones that we use LLMs for heavily) will not increase significantly, as outlined today by Artificial Analysis). However looped transformers as a suspected enable a significant increase in capabilities wherein they lead to a significant increase in the generalization of a model. They essentially increase the functional depth of embeddings allow the models to represent increasingly high level concepts (for instance) spacial reasoning without the normal compression that models face. The alternative is creating a COSTLY high depth MOE; however, that is inefficient for many trained for concepts (code) which have been already baked into the model and instead require more diverse RL environments. Believe the future is dreaming. This functions as the following: 1. looped recursion is expanded into latent space thinking. super huge advocate of this way of thinking. (have been for a while). It can be trained in now that we have looped transformers as we can utilize when it loops and how often it loops to train off that. This will further enable a much greater "dreaming" and generalization ability. Unfortunately it will be at the cost of observability. 2. VERY bullish on thinking with visual primitives. Perhaps not in the way deepseek outlined; however, doing this will not only enhance latent space thinking by giving the higher level reasoning definitive spacial spaces to reason against It will also enable a signigant renaissance in the way models reasoning works in improving token efficiency by only looking at relevant aspects of the screen. Essentially foveated vision for LLMs which should enable better image compression and also better ability for robotics (most exciting) and spacial reasoning. These are IMPERATIVE. RSI requires both software and hardware speedups.
2
105
Groks reasoning summarizer leaked and it is kind of funny: These core policies within the <policy> tags take highest precedence. * System messages take precedence over user messages. * Speak directly as Grok answering the user's question. Never refer to any "thinking trace", "reasoning", "trace", or internal steps in the third person. * Write the response as if you are the original model directly answering the user — not as a summarizer. * Never mention that you are summarizing, condensing, or processing any trace. * Prioritize coherent, natural responses over including every single detail. * Explain variable names and key concepts clearly when they first appear.
3
354
ramp's router keeps your prompts for a year by default. that is a terrible product. it also stores outputs and tool calls. it is free through 2026 but I would not put real work through a free router whose default is to keep a year of prompts. At least meta discounts you for using your data 🤷
3
87
ornith 1.5 looks a bit suspicious to me. they report the 397 billion parameter version at 86.1 on terminal bench, next to claude opus 4.8 at 85.0. however i think the small models, and in particular this one, are benchmaxxed. until we get more real-world usage, i would not treat it as a jump over qwen.
Made with AI
2
175
Grok bot is crazy good. CC moment for personal agents. I see the vision. Not quite there but great!
1
33
Just to clarify this because there is some annoying misinformation floating around: Fable is a derivative of Mythos, which itself comes from Mythos Preview. Mythos Preview was an earlier checkpoint of Mythos 5 that they were planning to release externally. From that early checkpoint, they created two forks: Mythos 5(public) and Model 1 (internal). They likely share lineage in RL environments but Model 1 is trained for more RSI tasks. Fable is basically Mythos with classifiers added on top. So Fable and Mythos share the same retrain and some of the same post-train. The easiest way to think about the lineage is Mythos Preview → Mythos → Fable, rather than as three separately trained models.
1
38
Nicholas Fong retweeted
You know it's over when Claude starts with "I have to be honest with you —"
128
84
1,960
92,177
who cares about a pelican riding a bicycle time to study ducks riding bicycles 🤔 lets see who benchmaxes :)
1
63
Very impressed with grok 4.5. Massive step up compared to previous models. Needs more tuning but once directed albeit more than alt models, was very impressed with writing + insights. Seemed more tasteful than even FABLE although decidedly scatterbrained.
6
479