cofounder & CMO @cloneisyou | love to learn new language 🔤 and food 🍱 | community: discord.gg/fqwAmDxa9

SF
Pinned Tweet
Frying chicken after event
1
1
15
1,729
1. This is the most thought-provoking piece I’ve read in a while. The author’s incisive analysis gets to the heart of what intelligence is. 2. How I came across it makes it all the more interesting: through a carousel posted by a Korean Instagram account, rather than on X itself. With Instagram’s format built so firmly around images, I was used to encountering lighter, more mainstream content there. I was surprised to find a real audience for a post like this. 3. Even on an internet that seems to be growing flatter, then, there is still plenty of room for translating value and for arbitrage. Once, there was plenty of value to be captured simply by translating Steve Jobs’s Stanford commencement address into Korean. Today, automatic translations let people watch videos of well-known figures directly, even when English isn’t their first language. 4. Translation is now needed for discovery, rather than comprehension. With the editorial discernment to spot a Chinese-language post on X and turn it into an Instagram carousel, you can still cultivate a real audience. A line from Anna Tsing’s The Mushroom at the End of the World, which I’m currently reading, has stayed with me: “Translations across sites of difference are capitalism: they make it possible for investors to accumulate wealth.” (p. 62) This may not be the most faithful use of a passage from a book that explores the possibilities of survival in pericapitalist worlds. Still, her words have become entwined with this post in my mind, and I wanted to put that connection into words.
在机器学习和信息论中,有一个著名的直觉叫流形假说。 现实世界产生的数据,在名义上往往拥有极高维度:一张高清图片可能对应数百万个像素维度,自然语言、声音、蛋白质序列、行为轨迹也都存在于极其庞大的状态空间中。 但有意义的数据并不会均匀散落在这些空间里,它会被压缩在一个极小、极其特殊的结构区域中。 一张随机生成的百万像素图片几乎不可能恰好构成一张真实人脸,因为自然图像同时受到身份、姿态、光照、透视、物体结构和物理规律等大量约束。 流形,只是这种低复杂度结构的一种几何表达,更一般地说,机器学习真正能够成立,是因为现实世界的数据根本不是随机噪声。 这也是理解智能最重要的起点。 智能存在的第一条件,在于这个世界必须具有可预测结构。 假如一个宇宙中的未来和过去完全独立,那么无论是人脑、Transformer还是无限大的超级计算机,都无法通过观察过去提高对未来的预测,因为这个世界根本没有留下任何可以学习的规律。 白噪声无法被理解,也无法被预测。 因此,智能能够存在,本身就意味着宇宙不是最大随机性的混沌,而包含大量跨尺度稳定的统计依赖、几何约束、组合结构、动力学规律和因果关系。 从这个角度看,机器学习做的事情是在巨量观测中寻找可以压缩的规律。 模型看过十亿张人脸,并不需要在参数里存储十亿份像素副本,它需要形成的是关于身份、轮廓、姿态、光照、空间关系等潜在变量的生成结构,最终用一个远短于全部训练数据本身的模型去近似。 因此,Learning在很深的意义上就是Compression。 如果一个庞大的数据集合能够被一个更短的生成程序描述,那么模型获得的就是规律。这与Kolmogorov Complexity的直觉高度一致。 生成或描述这些数据所需的最短程序究竟有多长,是一个重要的问题。 但是,数据不均匀本身仍然不足以产生智能。一个变量即使99%的概率取0、1%的概率取1,它的分布极度不均匀,却未必具有任何复杂结构。 重要的是变量之间存在可以利用的依赖。知道 (X) 之后,是否能够减少对 (Y) 的不确定性,这才是信息。 Mutual Information (I(X;Y)) 所度量的,本质上正是这种关系。 所以,智能依赖的是世界内部存在稳定的dependency; 统计规律、几何结构、层级结构、因果关系以及时间上的连续性共同构成了学习能够发生的基础。 这也给出了理解大模型的一种更准确方式。 一个随机初始化的神经网络最初只是一个巨大的、没有意义的函数族,而梯度下降不断利用数据排除那些与现实分布不一致的函数。 经过海量训练之后,原本随机的参数逐渐形成高度结构化的内部表征。 语言结构、空间关系、人物与概念关系、程序结构和部分物理规律也进入其中。 这里,我们不能简单说神经网络与现实世界同构。因为同构是一个过于严格的数学概念。 其实,模型是在自己的计算介质中形成了一套能够保存现实世界部分稳定关系的内部表示。 它是世界结构在另一种介质中的压缩投影。 于是,所谓智能涌现也没有那么神秘。 如果模型只能捕捉局部统计相关性时,它表现为模式匹配; 当它能够形成更长程、更抽象、更稳定的表示时,就出现泛化; 当不同领域的内部结构能够被重新组合,就表现为类比和联想; 当模型能够利用已有结构,对训练数据中没有直接出现过的状态进行推断,就表现为推理。 记忆、泛化、类比和推理,只是同一种机制在不同尺度上的表现:利用一个内部世界模型,在未观测区域重建结构。 模型规模之所以重要,是因为复杂世界需要足够大的函数空间才能被表示。 而某些能力的突然涌现,有时只是评价指标造成的阈值效应,有时也可能对应grokking或内部表示重组这样的真实非线性变化。 我想,我们需要更多地去关注参数空间在学习过程中发生了什么结构重组,使新的计算能力成为可能。 如果继续把这个问题推到信息论和统计物理层面,会得到更加统一的图景。 生命,依靠持续消耗自由能维持自身远离热平衡,而大脑和机器同样需要消耗能量进行信息处理。 由于Shannon entropy和热力学熵并不是同一个概念,我们可以更准确地定义,智能系统不断丢弃与任务无关的变化,同时保存对于预测和行动有价值的信息。 这与Information Bottleneck极其接近。 内部表示并不需要保存输入的一切,而只需要保留足够多能够预测未来、指导行为的结构。高级智能的标志,是知道什么必须留下,什么可以被遗忘。 由此,智能可以粗略地被写成Compression、Prediction、Control。 Compression意味着从海量观察中提取规律,Prediction意味着利用这些规律推断尚未看到的世界,Control意味着利用预测去改变未来。 一个有限的系统面对一个远比自身复杂的宇宙,通过观察世界、压缩世界、建立内部模型,再依靠这个模型预测并干预世界,这就是智能最一般的形式。
2
6
355
1. Outsourcing execution 2. Accelerating execution With these two shifts brought about by agents, I feel we are moving toward a way of working that makes it harder and harder to stay absorbed in a single task for long. Losing that sense of immersion is frustrating in itself, and it also makes us less effective at our work. To get the most out of agents, you inevitably end up running many sessions in parallel. Yet, as Garry pointed out in a recent YC talk, our working memory has room for only about the “magic number seven.” Juggling different kinds of work, each complex and laden with context, is an ordeal. I feel like a seven-digit computer suffering a cache miss on every read. Most of my “bio tokens” now go toward reading and making sense of what I have read, because agents handle most of the writing and execution. Reading, meanwhile, demands an “Extra High” level of mental alertness and focus, so the number of hours I can devote to it each day is more limited than I would expect. Lately, I cannot shake the feeling that, despite having the ability to put agents whose knowledge far exceeds my own to work in parallel, my biological limits keep me from making full use of them. How can I create the conditions for deep, sustained focus again? I have been trying what I can at a personal level. I tend to eat smaller meals, for instance, because I find it hard to concentrate unless I am a little hungry. But I wonder how long I can keep a clear enough view of everything to manage what feels like running an entire city. And as we run a race toward a destination already set, I wonder whether we really have the luxury of choosing between speed and the scenery.
2
7
266
Jun retweeted
Anthropic’s Economics team is sharing a new model of how AI might affect economic growth, jobs, wages, and more by 2030. Explore the scenarios, tell us what you think will happen, and see how your answers compare to more than 10,000 Americans. anthropic.com/institute/econ…
1,133
3,803
25,520
18,988,538
sf founders scraping by on seaweed gathered from the Pacific @cloneismin
3
562
Jun retweeted
got such a good vibe from @songwoojin korea x japan founders meetup in sf 🇰🇷🇯🇵
1
1
18
438
please expedite release date
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.
4
433
That's what I need exactly
ChatGPT can now remember your activity across the apps and websites on your computer. With Computer History in the desktop app, future interactions feel more personalized and require less explanation.
2
187
Jun retweeted
Google Brain founder, Andrew Ng: "Prompting will die in 6 months. Loops and Graphs are what's replacing it." In 2 hours, he shows exactly what the best engineers already build instead, and how to start building it yourself. The missing piece most people skip: how to connect those loops into a graph that compounds every time it runs. Watch it, then read the full guide on loops and graphs below.
82
738
4,608
1,048,352
At Salesforce park
5
217
Please say hello to me
1
2
156
At Corgi Cafe I love their wifi password: getinsurednow
2
7
234
Ohhh thank you @nico_laqua!
2
98
OpenAI Codex Community Meetup We met great independent researcher @barisozmen_twi I also hope AI lets us become Gentleman scientist 😏🤓
1
9
374
So, can I get it like, OpenAI's pace is slow enough, right?
4
81
Open to work: Chief Food Officer I cooked all of this. - N years of Korean BBQ experience - Galbi marination + full-stack grilling - 1 month as an Army cook, feeding 150 people at scale - 0-to-BBQ execution Startups hiring a CFO, DMs open.
Craving Korean BBQ in SF? Come hang out at Clone House SF. Founders, developers, and researchers welcome.
3
116
Interesting research question. Excited to see where this line of work goes!
Excited to share my first work after joining @MSFTResearch! LLMs have entered the agentic era, and we now collaborate with agents on complex tasks over many turns of interaction. But does your agent actually follow what you intended? We show where agents get lost: LLMs Get Lost in Evolving User Intent. 📄 arxiv.org/abs/2607.20734 🧵 1/N
4
284
The hardest part of education isn’t teaching, it’s evaluation. Once a metric becomes a target, people optimize for the metric, not the underlying ability. AI makes learning easier, but it also makes reward hacking easier. To change education, we have to redesign what gets rewarded.
3
67
Jun retweeted
Cheif Food Officer Cheif Dance Officer ABCDEFZ @hyeongjun_12185
1
1
8
308