Product assurance AI agents that use your web product like real users. Signup, checkout, dashboards. Catch UX bugs before merge. muggletest.com

Pinned Tweet
I vibecode even when I am AFK too. You are one command away from this smooth mobile vibecoding experience. Try Muggle Test’s open-source agent harness FREE: Ask me anything in the comment below. /plugin marketplace add github.com/multiplex-ai/mugg… /plugin install muggleai@muggle-works
1
4
291
Stan个人观点: “现在LLM-based的大模型本质上是”依赖人类去做剪枝的暴力穷举解法”,通过天量的参数量和注意力架构优化,加上overfit人类的推理习惯数据,压缩特征知识去模仿接近人类的联想思考过程,真正有价值的地方在于提供一种新的知识抽象和压缩解压的算法方案” “可惜人类的思考不是线性的,context window也不是这么work的,attention也不是单线程的,通过工程手段去补缺漏并不能弥补基础智能上的缺陷,我认为离AGI还差得至少一代技术革命” 你认为呢?
1
1
1
25
Muggle AI retweeted
有朋友问,为什么教皇坚持说AI没有意识,Anthropic却说有可能有? 确实有这个分歧。9月29日,《纽约时报》报道,今年5月教皇良十四世发布AI通谕《Magnifica Humanitas》之前,Anthropic的联合创始人Chris Olah看到预发稿,一度想退出发布会,之后几个月里,Anthropic私下游说教廷的顾问和宗教学者,希望他们认真对待机器可能有意识这件事。 10月2日,教皇在X上说,艺术和机器靠统计算出来的东西之间,先有本体上的差别,然后才是审美上的差别,算法缺少人的火花。通谕第99节说得更直接,所谓人工智能,不经历什么,不感受快乐和痛苦,没有道德良知。而Anthropic给Claude写的宪章,对Claude现在或将来是否可能有某种意识、某种道德地位,表示不确定。 平常写这个系列,我尽量不选边站。但这一篇我不得不选,而且选得很明确,因为我自己的论文就是这么写的:今天的AI没有意识。在这件事上,我站在教皇这一边,只是理由跟他不一样。 教皇的理由,是人有身体,有灵魂,有在关系里长大的一生,机器没有。结论我同意,可这条理由建立在一整套神学的人观上,不信这一套的人不会被说服,它也说不出什么样的证据会让它改变看法。 我的理由不靠神学。我在《人工智能意识的不可能性定理》里的结论是,纯确定性系统不可能产生意识,这不是当前做不到,是结构上不可能。推理是一条链。意识的核心是否定,是判断和拒绝的能力;否定要有余项,也就是在时间里积累起来、任何规则都收不尽的那部分自由度;而余项的必要条件,是真随机乘以结构化时间。真随机,指的是物理层面的不确定,不是电脑里的伪随机数;结构化时间,指的是这些随机在时间里被保留、被筛选、被积累下来,生命的演化就是一种。这是乘法,任何一项为零,乘积就是零。 今天的AI,两项都是零。它的计算是确定的,采样时用到的那点随机,是伪随机,给同一个种子就能原样重来。它的形态也不是自己在时间里长出来的,是训练从外面灌进去的,我把这叫注入,注入跳过了真随机在时间里的积累和筛选,时间这一项也就是零。 所以,当一个模型说“我害怕被关掉”,或者说“不”,那是在表演否定,不是在执行否定。这些话都能追溯到训练数据里的模式。Suleyman批评Anthropic循环论证,先在训练里让模型琢磨自己可能有意识,再把模型的话当成证据,他说得对,而且问题比循环论证更深,这些话从结构上就当不了证据。真正的否定,得是一个系统否定了自己的优化框架,而且这种否定追溯不到训练数据,这比说一个“不”字难得多。我在死后那篇写过,人跟桌子、杯子的区别,在于人会说不。今天的AI也会说“不”,只是那个“不”,是别人替它说的。 朋友追问说那Anthropic的“也许”算不算一种谦虚?我理解它的谨慎,可不确定不等于五五开。结构上已经排除的事情,还说也许,就不一定是谦虚了,是把一个产品的表演和一个主体混在一起。尤其当这家公司一边卖这个产品,一边请神学家吃饭,希望教会认真对待它可能有意识,这个“也许”就更难让人只当成哲学。 当然还要说清楚一条边界,我跟教皇的区别也在这里。我没有说机器永远不可能有意识。如果将来某种人造系统,真的引入了物理层面的真随机,又有自己在时间里积累的过程,意识在理论上就不被排除,但那时它已经不是今天这种AI了,而且这只是必要条件,不是充分条件。教皇的门是用教义关上的,我的门是用条件关上的,条件可以检验,也可以被证伪。 教皇说AI没有意识,是因为它没有灵魂;我说它没有意识,是因为它没有自己的时间。
20
4
33
10,148
Ride to Rebuild is finished. 21 Foster City households rode 2,390.6 miles in September. $2,578.90 is going to the families displaced by the Sea Spray Lane fire, and it's already with the fund. Thank you to everyone who rode. muggle-ai.com/charity
1
1
2
13
个人观点: 现在LLM-based的大模型本质上是”依赖人类去做剪枝的暴力穷举解法”,通过天量的参数量和注意力架构优化,加上overfit人类的推理习惯数据,压缩特征知识去模仿接近人类的联想思考过程,真正有价值的地方在于提供一种新的知识抽象和压缩有损解压的算法方案 可惜人类的思考不是线性的,context window也不是这么work的,attention也不是单线程的,通过工程手段去补缺漏并不能弥补基础智能上的缺陷,我认为离AGI还差得至少一代技术革命 你认为呢?
1
20
Today is the last day to ride for Ride to Rebuild. 2,297.5 miles by 20 Foster City households so far — $2,485.80 unlocked for the Spinnaker Cove Fire Relief Fund. $514.20 of the $3,000 still unclaimed. $1 an adult mile. $10 a child's ride. muggle-ai.com/charity
1
1
11
this is lit 🔥🔥🔥🔥🔥🔥🔥🤟🤟
I still can't resist laughing, watching this on LooP again and again and I will pass out laughing 😂😂😂😂 I love the beat though, and WTF is Sam upto, just watchin him 🤣
1
1
22
Way to go brother!
EverOS is now a @convex-dev component. 🧠 Add persistent long-term memory to any Convex app, or to @convex-dev/agent in one line, backed by EverOS Cloud. Convex already gives agents the backend primitives they need: a reactive database, durable functions, scheduling, and threads. Now agents can remember across threads and sessions too. EverOS turns conversations into durable user memory: facts with timestamps and traceable sources, plus a profile that evolves over time. That memory belongs to the user and follows them across threads, sessions, models, and agents. Under the hood, EverOS delivers state-of-the-art results across long-term memory benchmarks. No vector database to run. No memory infrastructure to stitch together. Get it today: convex.dev/components/everos…
1
4
25
AI agent feels like one of those "underpaid employees" - unincentivized, has no interest in doing a better job, do the bare minimum, forgetting, passive, and make mistakes all the time. No way this is AGI LOL Comment below what you think
1
1
12
Too old to be invested😥😥😥
最近真的有点看不懂这个创业市场了。 仿佛现在00后创业都不够年轻了,得是03后、05后,最好再加个“天才少年少女”的标签,才能成为资本市场的新宠。 可问题是,最大的00后也才26岁啊! 更魔幻的是,我身边甚至有95年的朋友,为了迎合这种风气,在某书上把自己包装成03年的AI连续创业者。不可否认,他本身是一个很优秀的人,但我实在想不明白,为什么一个优秀的创始人,还需要靠虚报年龄来证明自己的价值? 我一直认为,年轻确实意味着更多可能性,但创业终究是一场长期的考验。过往的创业研究也表明,成功创业者并不都是二十出头的年轻人,很多人在35岁甚至更晚才迎来事业的突破。 不知道从什么时候开始,我们似乎越来越热衷于制造“天才少年少女”的神话。今天把一个年轻人捧上神坛,明天又可能因为他没有达到外界的预期,而迫不及待地将他拉下来。 我不觉得这是一种适合年轻创业者成长的环境。 年轻本来是一件很美好的事情,但不应该成为一种需要不断刷新下限的竞争指标。一个人的价值,也不应该因为他是95后、00后还是05后而被重新定义。 比起一个又一个关于年龄的造神故事,我更希望看到的是,资本市场能够给予年轻人真正的耐心,让他们有时间试错、成长,慢慢做出真正有价值的东西。 毕竟,创业不是一场比谁出生得更晚的比赛。
1
2
21
“Giving an AK to a money and see what will happen” is never a good idea. Exactly what the top AI model companies are doing.
1
1
15
🙌🙌🙌 AI makes work cheap. It doesn't make it right. Focus on the people
Agents can be AI, but agency is human. It is every individual’s responsibility to empower our own agency, with the help of tools. But the North Star should always remain human centered.
1
1
18
IMPORTANT: If you are using Claude Works: SAVE EVERYTHING ON YOUR DISK!!!!! Plans, artifacts, decks, data, everything. No idea why people so hyped about Claude Works. It is such a sub-par product. Here is my experience: I tried oneshotting my sharing's deck with some ideas. The content it created is just...bad. And I ended up spending another day to handhold the agent, slide by slide, point by point, sentence by sentence. No way it can be more frustrating than this, right? ...The chat session disappeared after one day. No where to be found. Not phone client, not desktop client not even web client. Not in projects, not in code, not in chat. Just a lot of PR.
1
1
29
Muggle AI retweeted
I am hosting a closed-door sharing session for a few audiences about my first-hand insights on LLM/AI technologies as an AI startup founder this Friday (09/18/2026) If you would like to join, DM me.
1
3
36
See how Muggle Test testing a new navigation bar feature for "Ride to Rebuild" campaign! ON MY DEV STATION!! Try us: muggletest.com/
2
2
19
Muggle AI retweeted
我发现,几乎所有想靠 AI 赚钱的人,最后都会经历这五个阶段。 第一阶段:钻研技术。天天研究新模型、新工具、插件,skill,越学越兴奋。最容易产生一种幻觉:我懂了这么多别人不懂的东西,肯定能赚到钱。 于是进入第二阶段。 第二阶段:搞“一人公司”。觉得自己会 AI,就不应该再用传统方式做生意。运营、设计、程序员、客服,全部让 AI 干。一个人、一台电脑,顶一个团队。结果做了一阵才发现:Token 越烧越多,订阅费越来越高,钱却没赚到多少。这时候才明白:AI 只能提高效率,但如果你做的东西没人要,效率越高,只是更快地生产垃圾。于是进入第三阶段。 第三阶段:开始搞流量。你发现,产品再好,没人看见也没用。于是开始研究自媒体、标题、短视频、算法、完播率、互动率、涨粉。以前研究的是:“AI 怎么用?”后来研究的是:“怎么让更多人看到我?”然后你会发现,AI 变现翻来覆去就那么几种:卖课、咨询、社群、定制开发、卖资料。于是进入第四阶段。 第四阶段:开始教别人赚钱。把自己踩过的坑、研究过的方法打包成产品。开始讲:“普通人怎么靠 AI 赚钱。”“我是怎么用 AI 涨粉的。”“一个人怎么靠 AI 做公司。”结果又发现,教人赚钱这条赛道早就挤满了。用户也被各种“副业”“搞钱”“认知课”割麻了。于是进入第五阶段。 第五阶段: 就是彻底佛系了,对AI去魅了。 极少人进入第六阶段:找到一个真正适合自己、有人愿意付钱,而且可以长期做的方向。把技术、流量、产品和商业模式真正串起来。
281
197
1,447
258,826
+1
Dario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead. You guys are the frontier. By any reasonable metric — market share, revenue growth, model capability — the two of you have a duopoly on frontier intelligence. You’ve also claimed the lead is widening because of recursive self-improvement. I don’t see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible. But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff. Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier. Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want. Pacing the frontier would also create breathing room for a more intelligent conversation about regulation than Bernie Sanders’ “shut it all down.” China is very unlikely to join a global agreement, as you know, and that has to be taken into account as well. So go ahead and pace the frontier. You are the ones setting it. The easiest way not to build superintelligence is for you to agree not to build it. Demanding your preferred regulatory framework as the price of that will look like blackmail of the public and the political system. So just do it. If you do, you’ll buy goodwill for the next conversation. If you don’t, we’ll know this was just another bid for regulatory capture — or an election-season psyop.
1
23
We're giving out 10 Muggle AI x MIIR coffee mugs to our early users! If you haven't already, register your account at muggletest.com. Comment and repost this. We'll randomly draw and contact you for your shipping address by 10/01. Try us: muggletest.com
25
Late night random thoughts as a founder: AI agent prob not the answer, just hype: it think and work linearly, forgetting, gives you unpredictable errors, verbose communication without a clear point. Feels like a sloppy assistant Sometimes I feel what did I do wrong to deserve all these pain
1
15