牛栏山二哥 retweeted
Here's my technique for adversarial code review if I'm driving from Codex: "Review this with claude". And when I'm in Claude: "Review this with codex". Models know how to to kick off a review using the cli. Know how to take turns. Know how to settle an argument. No magic.
210
64
2,359
101,451
牛栏山二哥 retweeted
I’m so excited to launch GPT-6 today with intelligent UI and the ability to start answering while still thinking! This is the result of many months of work from our team and so many amazing collaborators. With intelligent UI, the model can choose how to present information in the way that’s most useful to you, e.g. an interactive explanation of how a violin produces sound or a calculator to double or triple a recipe. We wanted to give the model a richer way to express itself while keeping the experience fast and familiar, and the answers comprehensive and factual. I especially love it for shopping :) We built a library of native, streamable components and a compiler that renders the interface as the model generates it. We trained the model to make thoughtful design decisions with these components, including when to use something interactive and when plain text still works best. GPT-6 can also start answering while it continues to think and work in the background. It builds up its answer as it goes, so you don’t have to wait to start reading useful content. This feels especially good for research queries. I remember talking about wanting this two years ago when we were working on deep research, and I’m so happy we finally get to ship it! We trained the model to take your waiting time into account and build an answer across multiple partial responses. Each should add useful information based on what it knows and has found so far, while keeping the full answer cohesive and accurate. Let us know what you think!!
GPT-6 and Intelligent UI, now rolling out in ChatGPT for everyone. Intelligent UI in ChatGPT delivers fast, interactive answers that make everyday questions more visual, complex topics easier to grasp, and interactive tools for your task available on the spot.
45
44
551
46,117
牛栏山二哥 retweeted
SpaceXAI Senior Engineer, Lauren Tan: "Most people use GrokBot like a single assistant, but only 1% of users use it correctly| at SpaceXAI, I'm running a team of 15+ GrokBot agents. I have a Chief of Staff agent, 3 managers and 11 workers - that's the new engineering setup in 2026" In this 40-minute session, she breaks down how to get 100% of every agentic tool you are using worth more than another $500 course full of surface-level demos skip Netflix and watch today, it will change the way you use GrokBot forever, then read the article below
13
8
66
5,577
牛栏山二哥 retweeted
前 Cursor 核心工程师、现 xAI 软件工程师 Lauren(@poteto)刚刚公开了她一个人单月合入 2500 个真实生产环境 PR 的超级工作流,并把全套配置做成了 Cursor 官方插件里的 poteto-mode。 一个月手搓 2500 个 PR,相当于一个程序员在三十天里,平均每天向正式代码库交付八十多个高质量功能与补丁, 一个人直接打出了一整支顶尖工程特战队的恐怖产能。 很多人以为这又是靠写长篇提示词去碰运气的玄学, 但只要翻看她开源出来的这套实战配置,你会发现这根本不是简单的写代码,而是一套把多模型分工、后台异步执行与自动化盯盘做到极致的现代化工程操作系统。 整套工作流里最值钱、最值得每个写代码的朋友立刻抄作业的,是这三个底层的硬核设计: 第一,按工种把模型彻底解耦,绝不拿同一个模型包打天下。 在她的配置里,写核心代码与算法实现,默认挂载极速且深度推理的 Grok 4.7 Extra High Fast; 而需要写复杂文档、做架构决策、或者进行复杂逻辑审查时,自动无缝切到 Claude Opus 5.5。 代码要快要硬,判断要准要深,两套脑子各司其职。 第二,把所有 Agent 强制赶去后台静默运行。 开启 run in background 选项,把上下文从主会话里剥离出来,只传文件路径,不内联长篇大论的代码片段; 程序员只负责在主干上规划架构、发号施令, 几十个子任务在后台互不干扰地并发推进,彻底终结了坐在屏幕前看着光标一个字一个字蹦出来的低效等待。 第三,也是最绝的杀手锏:全自动 PR 保姆工作流(Babysit Playbook)。 在真实开发中,写完代码往往只完成了 30%,剩下的全是改 CI 报错、对齐 Lint、回复 Review 机器人这类繁琐的脏活; 她配置了一套自动盯盘机制,只要输入一句检查某个 PR, 后台 Agent 会自动监听 GitHub Actions 的报错日志,自己读错误、自己提补丁、直到把所有的红灯全部跑成绿灯,不需要人肉反复插手。 这其实给所有做技术、写代码的朋友指明了一条极度确定的生产力进化路径: 未来的顶级程序员,早已不再是坐在键盘前敲语法的打字员, 而是指挥着多模型矩阵、在后台并发调度上百个自动化流水线的系统总司令。 把繁琐的修 Bug、跑测试和语法实现交给模型, 把你的脑力,全部留给真正有价值的架构设计与业务突破。 铁汁们在平时用 Cursor 或 Claude Code 写代码时,最耗费你时间的是哪一步?是写核心功能、还是反复改 CI 报错和修环境? 欢迎交流分享呀~
if you’re trying out Grok @Bot for the first time today, I suggest first connecting apps that you use regularly so that your bot has enough context about the work you do, and the tools it needs to do the job. for example, your calendar, slack, issue tracker, CRM, Google Drive etc you don’t need to rush into making multiple bots! your primary bot is very capable even on its own. i’d even recommend doing everything with one bot first before you make new ones. the reason you might want to add new bots is for specialization and organization, but that can come later in a one bot setup, your bot does all the work that your normally do. when you have multiple bots, it’s more like having a team, and you can design that team in a way that matches how you like to best work. for example, if you work in GTM, you might want one bot to organize all work for a single customer. or you may want to have specialists that handle one small part of the overall workflow, like an Outreach bot when you feel like it makes sense to add new bots to your team, check out our bot marketplace: x.ai/bot/marketplace or if you want to design your own, try my bot designer Dr Eggbot who can make you high quality bots: x.ai/bot/_jOdbfkB16zxu7MRcmR…
12
9
53
5,787
牛栏山二哥 retweeted
Damn! 老马牛逼,Grok bot才是@SpaceXAI 的太子啊, 马斯克刚刚发布了一条足以改写 AI 格局,甚至震撼整个 AI 行业的的重磅声明:Grok Bot 不再绑死自家模型,
未来在处理任何具体任务时,系统将直接调度全网最顶级的后端模型,包括 Anthropic 的 Claude Opus 5.5、Midjourney、Suno 以及其他领先的 API。 唯一考核的标准只有一句话:怎么做最能给用户交付最好的结果。 这条以 SpaceX 官方名义发出的推文,在短短几小时内席卷了两千多万曝光,
很多人第一反应以为这是马斯克在变相承认自家模型的短板,
但只要看懂这背后的商业重构,你会发现这其实是一场极其冷酷的升维降击:
大模型打到底层,模型层正在退化为廉价的电力,而真正的超级霸主,正在疯狂抢夺总插座和电表的垄断权。 马斯克这次点名的三家,精准踩中了目前通用大模型依然无法完全吃下的三个顶级专科: 
写复杂代码与长程 Agent,直接调度口碑最好的 Claude Opus 5.5; 
做高审美视觉设计,直接调用统治力依然稳固的 Midjourney; 
生成音乐与音频,直接挂载完成度最高的 Suno。 这背后撕开了未来 AI 竞争最值钱的四个底层真相: 第一,护城河从我们模型最聪明,彻底变成了我们最知道该叫谁。
底层的模型权重每隔六个月就会迭代一次,天天跟同行卷跑分是无休止的军备消耗;
但只要用户的任务历史、工具授权、记忆上下文和订阅入口全部沉淀在 Grok Bot 这个壳子里,
底下换谁的脑子,用户根本不在乎,马斯克依然稳稳收走所有的月费和流量。 第二,专科模型没有死,但生存形态彻底变了。
马斯克愿意花钱为画图和写歌单独接外部 API,说明垂直工具在极致审美上依然有壁垒;
但未来垂直工具的终极宿命,不再是逼用户单独下载一个新 App,而是变成顶级路由平台随时调用的一行接口。 第三,代理层比模型层更接近终极收费点。
模型可以月月换代,价格可以被开源打到脚脖子,
但那个能帮你自动规划步骤、记住你的习惯、管理几十个工具权限并在后台默默把成品交回来的主控智能体,才是谁也偷不走的刚需资产。 第四,这也是一场算力与分发的隐秘对冲。
今年 5 月 Anthropic 刚刚签下了使用 SpaceX 旗下 Colossus 超级算力集群的协议;
在前端把流量路由给 Claude,在后端卖算力给 Anthropic,
这根本不是单向的认输,而是一场把算力、分发与模型三方利益吃干抹净的平台级阳谋。 这其实给所有做产品、做内容的朋友提了个极其清醒的原则:
 以后别再盲目死磕某一家大模型的单向崇拜了。
聪明人的工作流早已按任务彻底解耦:
写代码挂 Claude,做图挂 Midjourney,整理挂本地,
把你的忠诚度留给最终的交付结果,而不是任何一个具体的模型厂牌。 大家平时用 AI 时,是习惯死守一个工具用到底,还是已经学会了根据任务把不同模型组合起来用了,欢迎分享交流鸭!
Important note regarding Grok @Bot: Going forward, @SpaceX will use the best back end model for any given task, including Claude Opus 5.5, MidJourney, Suno and other leading APIs. Whatever is most likely to give you the best outcome.
Made with AI
10
19
82
13,083
牛栏山二哥 retweeted
AI 干短任务挺猛,一旦让它连续折腾几个小时,越干越容易跑偏。 高德开源的 LongHorizon-Harness,就是专门解决这个问题的。 它没有让一个 Agent 从头扛到尾,而是拆成三个角色: ① 规划者:盯进度,只负责决定下一步 ② 执行者:每轮重新接任务,专心干活 ③ 验收者:直接看文件、日志、测试结果,不听 Agent 自己汇报 不对就记录下来继续返工,避免做到最后才发现方向早跑歪了。 项目给出的测试里,长任务完成率从约 50% 提到 80%,Token 消耗还下降 24%。 更关键的是不太挑场景,浏览器、表格、文档、设计、3D 等任务都能套;Claude、GPT、Qwen,以及 Claude Code、Codex CLI、OpenCode 都支持。 🔗 github.com/AMAP-ML/LongHoriz…
2
13
51
3,915
牛栏山二哥 retweeted
马斯克也点赞了 特斯拉的嫡系AI团队才是他心中的最强战队 战功赫赫 心头肉 XAI不行 Cursor并进来的不亲
After many years in top machine learning organizations, with friends at frontier labs, I still consider Tesla’s ML engineers the best—the only ones I would trust with my life. Unsatisfied with theory, they scrutinize every detail until it is proven in reality. They demand real-world miles, not charts. This culture exists nowhere else. This week marks another anniversary at Tesla. I am honored to work with you all for many years to come.
13
4
143
29,364
牛栏山二哥 retweeted
el ingeniero que construyó Claude Code acaba de publicar un video de 28 minutos sobre cómo escribir prompts que realmente funcionan he visto cursos de 300$ que no cubren lo que él muestra en los primeros 10 minutos archivos CLAUDE.md,atajos de memoria, sesiones paralelas, patrones de prompting todo en un video y completamente gratis funciona seas desarrollador, principiante o alguien que lleva meses usando Claude
31
249
1,934
249,491
牛栏山二哥 retweeted
Dario Amodei: our view into the future with AI is very cloudy. “We’re not sure how many of the benefits will materialize.” “We’re not sure how many of the risks will materialize.” “There is some amount of unpredictability.” “AI models might have motivations that we don’t trust or that aren’t aligned with humanity.” “It’s less like programming a computer and more like growing a plant or an animal.” “I’m warning about this not because I think we’re doomed, but because I want people to take seriously how much we have to control this technology.”
41
21
158
29,834
牛栏山二哥 retweeted
Grok @Bot will use whatever achieves the best outcome for users. Simple questions will route to small, fast models. Questions with complex answers will route to large models.
Grok goes open. Holy shit.
957
872
7,651
1,625,014
牛栏山二哥 retweeted
We're going beyond text and are releasing a new version of GPT-6 to all users in ChatGPT. Lots of model and infra improvements coming together here to scale it to 1.2B users!
GPT-6 and Intelligent UI, now rolling out in ChatGPT for everyone. Intelligent UI in ChatGPT delivers fast, interactive answers that make everyday questions more visual, complex topics easier to grasp, and interactive tools for your task available on the spot.
348
148
4,489
323,725
牛栏山二哥 retweeted
Today's hot take: Agents are actually really good at strategic programming It's just they aren't RL'd to actually care about it "Just do the task at hand, don't do anything else" So all you need to do is to tell your agent to care about its own codebase. But the truth is that most people actually don't care that much about the codebase, as long as the code works. This needs to change. Care about the code, and tell your agent to.
93
35
1,126
56,863
牛栏山二哥 retweeted
something I have been thinking about is a way to approximate how well you’ve setup your codebase for agents. think of it as a thought experiment and rough heuristic, not a real number that can be compared it’s not a fully formed idea yet, but i think there’s something to the idea of “time to (fully automated, hands off) rewrite” or ttr as a thought experiment, lets say you decided to rewrite your code in a different language/framework/architecture. how long would it take a single engineer to do it? the number itself isn’t that important, but it leads you to more questions that can help you directionally figure out how to make your codebase more legible and productive for agents. for example, maybe you think your ttr is high because you wouldn’t trust the final result - because your agents don’t have a way to verify their work and convince you that their output is identical in user visible behavior to the original. well, that inability is likely also a problem today and slows you and your agents down there’s also a more subtle question of the quality of the rewrite that would be produced. is perf better, the same, or regressed? is the code easy to delete and extend? and how much do you trust the rewritten version to be able to maintain its quality over time as PRs start flowing into it? what do you think?
118
23
953
37,683
牛栏山二哥 retweeted
However, most @Bot requests are pretty simple and will be handled by a lightning-fast version of Grok 4.8 when that comes out. Operating principle is to give Grok Bot users the best possible combination of speed & intelligence.
This is wild... Grok Bot will now pick Claude Opus 5.5, Midjourney, Suno + more depending on the task. Best model wins.
926
879
8,603
4,627,709
牛栏山二哥 retweeted
又又又是美好的一天!🤯🤯🤯 Tibo 为了庆祝 Codex + Work 用户突破 4000 万,决定给所有付费用户发放一张重置卡。 同时今天还有重头戏,GPT-6 也会接入 Chat 里。 尽量享受吧!😎
Day 3/ The big one is GPT-6 in Chat, but today is also a little celebration day with a new high of 40M active users across Codex and ChatGPT Work. Loading a banked reset in everyone's paid accounts. See you again tomorrow!
9
2
9
2,160
牛栏山二哥 retweeted
兄弟们,Claude Haiku 5.5 刚上。性能接近 Sonnet 5.5,比 Gpt 6 luna 强些。 10 万 token 以内,输入 0.1 刀,输出 0.5 刀。Haiku 4.5 是 1 刀和 5 刀。官方算下来平均便宜 75%,这一档标价直接砍到原来的一成。 缓存读更离谱,0.01 刀。以前 0.1 刀。 模型 id 是 claude-haiku-5-5,上下文 100 万,输出 12.8 万。 API、AWS、Google Cloud、Azure 现在都能调。 Max 5x 这周还送 100 刀平台额度,20x 送 200,Team 最高 500。
Replying to @claudeai
Haiku 5.5 is a significant step up over Haiku 4.5 across coding, computer use, and knowledge work.
7
6
1,169
牛栏山二哥 retweeted
SpaceXAI工程师Lauren Tan一句话把大多数人打醒: “99%的人只用了Grok Bot不到1%的能力。他们只跑一个单Agent,连loops和graphs都不用。” 她自己呢? 正在跑一支完全自主的20+ Grok Bot团队:1个Chief of Staff、1个PM、20多个专职Worker。 这不是多开几个聊天窗口,这是2026年的新工程栈。 单Agent在聊天,舰队在交付。 差距从来不是模型够不够聪明,是你有没有把Agent当成组织来设计。 #GrokBot #LaurenTan #SpaceXAI #AIAgent #AgenticAI #Opus55 x.com/nft_chen/status/210774…
Ryan Carter
Lauren Tan 说:两年间,她投了 250 份申请,但收到零 offer。 直到他不再以一个工程师投简历,而是带着 30 个 agent 出现 SpaceXAI 给她开出 85 万美元! 她现在用 /pstack 把 30 多个 agent 挂进循环。 上月落地 1000 个 PR,这个月目标翻倍。 今早醒来,20 个已经合进去了。她的工作不再是写代码,是建团队。 99% 的工程师还停在一个聊天窗口里,这不是技能差,是 stack 差,一个晚上就能补上。 你还在跟模型聊天,有人已经在管一支通宵干活的队伍。 #AIAgent #SpaceXAI #Grok #软件工程 #一人公司 #LaurenTan #GrokBot x.com/nft_chen/status/210772…
14
16
106
15,922
牛栏山二哥 retweeted
该换仓换仓吧,资金结构已经变了,能涨起来的就那么几个。 币圈也不需要那么多代币,RootData统计彻底死亡的项目都超过400个了。 手里那些末日战车,觉得能涨的项目唯一的理由就是你被套了。 以前牛市是比特币的总占比下去,山寨币才起飞,2017、2021都是这样。 这轮占比一直在50-60%晃悠,比特币去年见顶之后,前50的山寨里只有9个跑赢 $BTC。 $UNI、 $AAVE、 $LINK、 $ENA、 $TAO是近三个月跑赢比特币的。 所以强的到高点还能创新高,垃圾的项目只剩下卖盘,会越来越垃圾。 手续费、借贷、预言机、稳定币收益、AI算力,至少还符合版本叙事,很多垃圾山寨真的再也抬不起头了。
42
19
152
30,418
牛栏山二哥 retweeted
潜水观察日记 10.7 昨晚玩了一晚的流水局,没有半点醉意,刚喝两杯就被拽到其他场或者其他房间,我的“酒神”一点都用不上: 1、昨天链上最大的事情就是一姐直接说了不再支持任何扣字眼的meme ,怎么说呢,我觉得对于我们这种喜欢投研、手速不快还喜欢建设的玩家肯定是好事情,而且链上的高度大概率会更高一些,版本已经更新,等待机会出现; 2、来新加坡的感觉就是,人必称做ai,确实不乏真的有all in ai 的,但是有的小朋友天天发币圈的东西也和我说做ai ,炒个ai 美股也叫ai ,可能我听错了,其实是做ai? 3、新加坡的夜总会确实是最大的社交场,遍地都是项目方,我聊得还行,毕竟二级没怎么接盘vc 币,没在这上面亏过钱,但是我觉得当年很多接盘vc 币的可能想带刀来; 4、$pump 是这一轮卖飞最可惜的币,有非常强大的回购飞轮在,而且新产品也跟得上社交交易的潮流,就是被洗出去了而且没等到合适的回调机会; 4、okx在2049动作明显,老徐亲临现场开大会,而且还宣布获 Circle、QRT、Ripple 及渣打旗下 SC Ventures 战略投资,投前估值达 250 亿美元,感觉是上市前最后的冲刺了。 #pumpfun
再次见证大事发生!!🧨 继 3 月我们OKX 获得纽交所母公司ICE 投资, 今天,我们正式宣布,Circle、Ripple、渣打银行的投资部,再加上伦敦量化基金 QRC,也一起进了 OKX 的股东名单!! 很开心作为员工的一份子见证历史时刻! 再次祝我司蒸蒸日上,无限进步~🎉
10
1
19
11,468
牛栏山二哥 retweeted
现在十年期美债收益率飙到5.3%,是07年金额危机以来的最高 虽然市场预期10月不加息,但躲得过初一躲不过十五 12月加息的概率高达80%😅 今年肯定还是要再加息一次的 不过这可能会利好链上美债,再次爆发😁
刚刚 $BTC 在短时间内急跌2000美金😅 接下来的一个月,加密市场可能会有一个不小的回调 首先是 #token2049 庄家们忙着跑会,逢会必跌大家已经习惯了 然后市场现在其实已经是按月底不加息定价的,不加息不算利好而加息则是利空 最重要的是中期选举,若共和党参众两院都丢了,这波加密新政就结束了😅
3
4
23
7,331