so say we all

Shanghai China
Liang Liu retweeted
WOW, @vercel 亲自给 GLM 5.3 打广告了。 Vercel Developers 宣布,GLM 5.3 即将登陆 AI Gateway,并且直接给了一个相当狠的评价: DeepsecBench 得分最高的开源模型,成本只有部分同分闭源模型的 1/3。 从 Vercel 放出来的成本-性能图也能看出来,GLM 5.3 已经站到了开源模型的第一梯队,而且位置相当靠近整个 Pareto 前沿。 这其实也是 GLM 5.3 最近最值得关注的地方: 参数规模没有继续无脑往上堆,但靠后训练和 Agent 能力,把实际任务表现硬生生拉到了顶级闭源模型附近。 @Zai_org 后训练仙人,又多了一张权威的第三方认证。
GLM 5.3 is coming soon to AI Gateway. The highest-scoring open model in DeepsecBench, at 1/3 of the cost of some proprietary models with similar scores.
25
6
75
18,006
这一个上午,我一个字脏话都没讲喷过,血压也低了,还连连 “Good,Nice,Perfect,Bravo,你办事我放心!” 从未有过的 vibe coding 体验, #GLM-5.3 仿佛让我看到了 Dario 和 Anthropic 头上发生的🍄☁️💥 。 加油国产模型
1
42
A Go CLI for modern LLMs, supporting OpenAI, Azure, Perplexity, LLaMA, etc. Features include streaming, chat, prompt files, multimedia I/O, MCP tool calls, and an experimental agent mode for multi-step tasks with safety and budget controls. #golang github.com/kardolus/chatgpt-…
3
37
3,451
MIT just released a 700-page book that actually teaches machines how to think. Pdf: algorithmsbook.com/files/dm.…
59
831
5,349
380,105
小白条
1
34
新竿子开光 奶翘
2
38
The gpt-oss models from OpenAI are a synthesis of ideas from prior research. Here are 10 interesting papers that were directly used in gpt-oss… (1) Longformer: Introduces sliding window attention, a form of sparse attention that is utilized in alternating layers of both gpt-oss models. (2) StreamingLLM: Describes the concept of attention sinks in large language models (LLMs)—these are tokens within a sequence that the model assigns high attention or weight to, simply because the softmax operation prevents the model from assigning attention to no tokens at all. (3) Off-by-one attention: Proposes a solution to attention sinks by allowing the attention mechanism to assign no attention to any token. This is achieved by adding a bias term of 1 to the denominator of the softmax operation within attention. In gpt-oss models, a similar approach is used, but the bias term is learned rather than fixed at 1. (4) Switch Transformer: Presents several ideas foundational to modern mixture-of-experts (MoE) based LLMs. It’s important to note that many other papers, in addition to Switch Transformer, have contributed to this field. (5) RMSNorm: A streamlined variant of layer normalization that is both more efficient and has fewer trainable parameters. Both gpt-oss models employ RMSNorm. (6) RoPE: Stands for Rotary Positional Encoding, a hybrid absolute/relative positional encoding method used by gpt-oss models. RoPE encodes absolute position using a rotation matrix and incorporates relative position information directly into the self-attention mechanism. (7) YaRN: A method for extending the context window in LLMs, which is adopted by gpt-oss models. YaRN works by adjusting the frequency basis used within RoPE and further training the LLM to handle longer contexts. (8) Flash Attention: Utilized by gpt-oss models, flash attention leverages system-level optimizations to significantly improve the computational and memory efficiency of the attention operation. (9) DeepSeek-R1: While the specific reasoning or reinforcement learning (RL) training strategies used by gpt-oss models are not fully detailed, the DeepSeek-R1 technical report offers a comprehensive overview of how RL training with verifiable rewards is implemented at scale. (10) Deliberative alignment: This is the safety training approach used by gpt-oss models, designed to teach the models how to reason through safety specifications and determine when it is appropriate to refuse a request.
7
88
405
28,458
The freshest AI/ML research of the week Our top 9 ▪️ Sotopia-RL: Reward Design for Social Intelligence ▪️ Agent Lightning: Train ANY AI Agents with RL ▪️ Exploitation Is All You Need... for Exploration ▪️ Learning to Reason for Factuality ▪️ VeOmni ▪️ Is Chain-of-Thought Reasoning of LLMs a Mirage? ▪️ Cognitive Loop via In-Situ Optimization ▪️ Sculptor ▪️ CoAct-1 ▪️ Tool-integrated Reinforcement Learning for Repo Deep Search ▪️ RL-PLUS ▪️ SEAgent ▪️ CRINN ▪️ Training Long-Context, Multi-Turn Software Engineering Agents with RL ▪️ Beyond the Trade-off: Self-Supervised RL for Reasoning Models' Instruction Following ▪️ CompassVerifier ▪️ Are We on the Right Way for Assessing Document Retrieval-Augmented Generation? ▪️ Are Today's LLMs Ready to Explain Well-Being Concepts? ▪️ VeriGUI ▪️ Trainable Dynamic Mask Sparse Attention ▪️ LeanK ▪️ Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models ▪️ On the Generalization of SFT ▪️ SitEmb-v1.5 ▪️ AttnTrace ▪️ LaTCoder ▪️ ChartCap 🧵
12
49
257
16,057
“机器会思考…”
22
Do you know who i am? 🕷️🕸️ #SpiderMan2PS5 #InsomGamesCommunity
2
16
190
43,526
Rust is secretly taking over chip development - piped.video/AwFU-CrIB8I
4
23
208
12,242
Weirdly I’ve seen people writing of Mbeumo, Cunha and Gyokeres even before they’ve put on a United shirt or signed the contract. Look Critique is part of the game , no problem. But I’m curious: on what exactly are people basing their verdict on Cunha, Mbeumo and even Gyokeres? Is it purely goals and assists (because quite obviously the numbers don’t lie), or just that they’re not your personal favourites, so they get thrown under the bus? Because from where I’m sitting, it feel a lot less like measured analysis and a lot more like a pram missing a few rattles because we didn’t sign your favourite player. I’ve seen top players come through this club. Not all of them looked world class on day one. Talent doesn’t always shout from the rooftops immediately sometimes it whispers, waits, and then roars when the moment’s right but Cunha and Mbeumo have shown serious quality so I’m interested to know why the pessimism?. How far they can go? That’s up in the air but dismissing them this early is short-sighted at best. Patience isn’t just a virtue in football it’s a necessity. Players need time to adapt, to understand their role, to gel with teammates, and to grow into the shirt they’re wearing. Some of the most iconic players we’ve ever seen were nearly written off before they found their rhythm, Judging them now, through the narrow lens before they’ve put pen to paper is I hate to say it stupidity. Let’s also not pretend your opinion is strictly tactical it reeks more of allegiance to someone else you’d rather us sign. Which is fine, but let’s not dress it up as footballing wisdom. Nothing in football is guaranteed. Momentum can turn overnight, systems can click, confidence can spark. So maybe hold off on sowing seeds of doubt in the minds of the easily influenced. Let these players breathe, develop, and prove themselves. You might just find they’re better than you gave them credit for. Some people have a platform with a picture of a legend act like Twitter Football Guru’s when in reality they’ll never sit face to face with you in a debate because their true allegiance to our club is as real as a WWE power slam.
15
7
80
4,100
端午节野钓翘嘴一条
2
131
3x Premier League Champion 1x Champions League Winner 1x FA Cup Winner 2x League Cup Winner 🏆
1
13
248
8,838
In Ruben, we trust.
77
848
12,720
318,882
fish at there fish are
32