AI analyst: models to products to money. Breakdowns on AI tools, funding and monetization for builders and investors. Signal over hype. EN + 中文. DM for collab.

SG
Grok 4.7 is now available on X and the grok website.
1
4
36
Nice customization attempt - should result in some really great mod repositories that further increase the playability.
You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:
1
56
对于流量,开始会有很多人告诉你 要做内容、做垂直,才可以 但是你点开他们的主页会发现 都是在说口水话、瞎扯淡,也有流量 这个时候你会发现原来这就是他们的“内容和垂直” 你猛然醒悟 “做内容、做垂直”原来就是接地气。
1
3
70
差不多时,必然考虑成本,但这种成本不应只是每百万的输入与输出成本,还需要包括所需的时间成本。
This week Claude Sonnet 5.5, GPT-6.1 Sol and Gemini 4 Argon all launched near the top of the Coding Agent Index leaderboard, but each has a different balance of performance and cost The Artificial Analysis Coding Agent Index measures agents (a combination of model and harness) across three agentic coding evaluations. ➤ Claude Sonnet 5.5 (max) in Claude Code takes the top spot at 68, but also has the highest measured cost per task: $14.19 ➤ Gemini 4 Argon (high) in Antigravity CLI scores 64 at $5.84 per task - less than half of Sonnet 5.5’s cost. Note, this uses Google’s promotional pricing, and Argon is not yet publicly available ➤ GPT-6.1 Sol (xhigh) in Codex scores 63 at $1.04, roughly one sixth of Argon’s cost
4
4
81
越是这样越让人觉得 Gemini 4Argon 是噱头,很难不认为是高级黑。
Current AI ranking after Gemini 4 Argon. 💀
2
4
206
若干年后,AI 和机器人统治世界,它们将会以此作为证据,一一追杀当初实施该计划的人。
Figure AI 让旗下的 Figure 02 人形机器人自己跳进芬兰一家铸造厂的钢水里,以此让这一代机器人退役。熔出来的金属会被加工成限量纪念品。 Figure 02 是 Figure 的第二代人形机器人,公司很多第一次都由它完成:第一次进 BMW 工厂上岗,第一次运行 Helix(Figure 自研的机器人 AI 模型,负责看画面、理解指令并控制动作),第一次做家务,第一次做物流工作。现在第三代 F.03 的数量越来越多,Figure 说继续维护 F.02 已经不划算。 退役的麻烦在于保护技术。机器人里有大量自研的执行器(可以理解为驱动关节的电机)等核心硬件,Figure 不想让它们流出去;但一台台拆解又太占工程师的时间,会推迟第四代 F.04 的发布。 Figure 于是在网上征求处理办法,施瓦辛格回复说"熔了它们"。他主演的《终结者2》结尾,T-800 机器人就是沉进钢水里的。Figure 联系了施瓦辛格,他表示想参与,项目随后启动。 找场地并不顺利。美国和墨西哥的铸造厂都不愿意让装着锂电池的机器人跳进自家昂贵的设备,Figure 还请《流言终结者》的前主持人帮忙找过。最后只有芬兰伊马特拉的一家铸造厂同意。 训练在 Figure 位于加州圣何塞的园区进行。团队铺了大气垫,让机器人从二楼往下跳。他们参考特技演员的动作,在仿真环境里训练了一个新的 AI 模型,让机器人在一个从没去过的铸造厂里,准确跳进盛钢水的容器。 铸造厂用的是一座 75 吨的电弧炉,靠三根石墨电极通电熔化废钢。团队只有 24 小时、6 炉的机会,每炉钢水只有 20 分钟可用,之后表面就会冷却结壳。Figure 称,现场的高温和强电磁场让摄像器材等电子设备出了故障,机器人则一直正常运行自己的 AI 控制程序。 熔出的金属锭已运回美国,正在加工成一批纪念 F.02 的限量藏品,数量很少,面向个人出售,博文没有公布价格和购买方式。大部分 F.02 已被熔掉,只有几台留在总部仓库。 Figure 在文末特意说明,视频不是 AI 生成的:机器人确实是自主起跳,被运到芬兰,跳进了钢水。
3
6
171
Would you trust a model that spent 17 hours writing a Lua compiler from scratch, debugging stack alignment and variadic returns, then passed 178/182 tests (97.8%)? Breaking:@AntLingAGI Ant Group's Ling-3.1-Flash did exactly that. 560B total params, only 25B active per token, 1M context. The architecture bet: 7 linear KDA layers + 1 Gated MLA per 8 layers, 512 routed experts with 8+1 shared. This isn't a benchmark story — it's a sustained-execution story. Model layer is becoming a cost-optimization lever. Whoever owns the serving layer owns the margin. #Ling-3.1-Flash
5
4
149
What if your Coding Agent wasn't a black box—but a plugin system where you install, write, and remix capabilities like LEGO? DeepSeek Harness v0.2 just shipped macOS/Win desktop apps. "Everything is a plugin": tools, UI, even the agent loop. 60% of users already run third-party plugins. New plugin manager: install by npm name, no CLI. Creative mode: describe a plugin, Harness writes it, persists instantly. Roadmap: sandbox hardening, agent teams, long-term memory, browser/GUI automation, mobile/remote, session sharing. Co-evolving with DeepSeek models—training optimizes for Harness, Harness innovates for models. 6¥ trial credit (1.2B cached tokens / 148K uncached / 520K output) for eligible users. My take: The winner in agents won't be the smartest model—it'll be the platform that makes extensibility feel like play, not plumbing.
DeepSeek Harness v0.2 端上官方桌面了! macOS/Windows 双包、开箱即用。「一切皆插件」:工具、UI、甚至 Agent 循环全是插件。60% 用户已经在跑第三方插件。新增插件管理页:输 npm 包名装、一键启停卸、无需 CLI。创造模式:对着说「帮我写个插件」,秒生成、即时生效、持久保存。 路线图:沙箱加固、Agent 团队、长期记忆、浏览器/GUI 自动化、移动端/远程、会话分享协作。与 DeepSeek 模型共同进化——训练为 Harness 优化,Harness 为模型创新。 6 元体验金(约 12 亿缓存 token / 148 万非缓存 / 52 万输出),合规用户登录即领。 #DeepSeek #DeepSeekHarness #Harness #崔添翼 #梁文峰 mp.weixin.qq.com/s/ZlHz-GhO2…
2
11
417
DeepSeek Harness v0.2 端上官方桌面了! macOS/Windows 双包、开箱即用。「一切皆插件」:工具、UI、甚至 Agent 循环全是插件。60% 用户已经在跑第三方插件。新增插件管理页:输 npm 包名装、一键启停卸、无需 CLI。创造模式:对着说「帮我写个插件」,秒生成、即时生效、持久保存。 路线图:沙箱加固、Agent 团队、长期记忆、浏览器/GUI 自动化、移动端/远程、会话分享协作。与 DeepSeek 模型共同进化——训练为 Harness 优化,Harness 为模型创新。 6 元体验金(约 12 亿缓存 token / 148 万非缓存 / 52 万输出),合规用户登录即领。 #DeepSeek #DeepSeekHarness #Harness #崔添翼 #梁文峰 mp.weixin.qq.com/s/ZlHz-GhO2…
3
4
640
正常来说, 发现有蓝朋友关注, 就应该回关回去, 这是基操。 但,有的人过了一天都不回, 你真以为别人是被你内容吸引了才关注你?
181
105
5,201
Anthropic just open-sourced the climbing gear. You're not "vibe coding" evals anymore. Anthropic just shipped /claude-api build-eval and /claude-api hillclimb—guided workflows that design evals mirroring production, split train/test to catch overfitting, and iterate one patch at a time until test set improves. The hillclimber reverts patches where train rises but test flatlines. It audits prompt rituals, steps down model tiers, and stops when gains drown in noise. Real numbers: support benchmark went from Opus 4.8 high-effort (74.4%, 4.6¢/ticket) → Sonnet 5 low-effort (98.9%, ~1¢) → held-out 90.5% vs 78.6% at 1/5 cost. claude-api skill: 66% → 88% by bucketing failures by root cause, not rewriting lines. The moat in AI apps isn't the model—it's the eval infrastructure that lets you climb the right hill without fooling yourself. #AIEvals #ClaudeCode
Claude can now help you build evaluations and hillclimb on them. In this article, we share guidance on eval design & skills that Claude Code can use to improve your applications. claude.dev/blog/automating-e…
2
222
Lisa Su just bought the only person who can tell AMD's engineers what "physical AI" actually needs from silicon. $8.2B all-stock. Fei-Fei Li → EVP & Chief Scientist, direct report to CEO. World Labs isn't a model company—it's a workload definition company. The chip war's next phase isn't who has more HBM; it's who defines the workload that justifies the HBM. Nvidia's moat: CUDA + ecosystem + partners renting their research. AMD's new moat: the professor who built ImageNet now draws the architecture diagrams for MI500. Markets priced AMD at $1T on "we sell shovels." This deal says "we now own the map." If physical AI (robotics, world models, sim) becomes the next scaling frontier, AMD just locked in the cartographer. My take: The highest ROI in AI chips isn't process node—it's knowing what to build before the market knows it needs it. #WorldLabs #AIChips #PhysicalAI #AMD
3
135
Would you rather prompt an AI—or hire, budget, and manage an agent team that ships work while you sleep? 91.9k ★. MIT. Self-hosted. One command: npx paperclipai onboard --yes. Paperclip gives you org chart, goals, tasks, budgets, agent templates—governance included. No vendor lock-in. The mental model shift: "I prompt an AI" → "I run a company of agents." OpenClaw is an employee; Paperclip is the company. Builders call it "Linear-level UX for agent orchestration." Eval infra just landed—Runner + Product E2E, public history hub. This isn't a wrapper. It's the operating system for the autonomous org. My take: The winner in agents won't be the smartest model—it'll be the platform that makes managing them feel like management, not plumbing. github.com/paperclipai/paper…
2
88
市场这种就应该做官方版本了啊,不是所有都应该让第三方做。难怪有人说是试验品啊。
今天提名几个广泛使用、内容丰富、更新及时的 DSH 插件导航站吧~ dshfind.com/ dshmarket.com/ awesome-dsh-plugin.com/ 其中 dshmarket 也可直接作为 DSH 插件安装到设置页中。 (不代表公司立场,不对非官方网站的内容负责。)
2
99
难道有人会拒绝用「同价、少说 7 成废话、分数不降」的模型? Fireworks 刚把推理模型的算账逻辑捅破了,Kimi K3 90% 的 token 都在「自言自语」。多轮 Agent 里,每一轮都把前几轮的推理塞回上下文——token 平方级涨,账单平方级涨。调低 reasoning effort?质量直接崩。Fireworks 的解法:后训练把「有用的自我纠错」留下,「无效循环」砍掉。Ember-1:Terminal-Bench 82% vs K3 Max 80.9%,DeepSWE 75.2% vs 66.4%,推理 token -35~50%。实战 A/B:单任务输出 29.9K vs 49.3K,推理 token -71%,总 token -39%,评分 0.753 vs 0.751。 但——只能走 Fireworks API,权重不开、训练细节不公开。这不是发模型,是推理平台垂直整合模型层、修自己的单位经济账。$15/M 输出,模型一闭嘴,实付 $9。 模型层正在变成推理平台的成本优化杠杆。谁拿着 Serving 入口,谁赚毛利。
1
57
还没有关注的蓝朋友在哪里? 推给我一些未关注的蓝朋友吧。 还是得互相浇水啊, 不浇水发再优质的帖子都没有流量。 #浇朋友 #蓝V互关
8
10
189