AGI forecaster. In 2022 I predicted an AGI timeline of 2027 MIT dropout

mei retweeted
Steven Sinofsky on why Jev fixes his oldest complaint about AI: natural language was never an efficient interface. "Ask yourself how many people are really, really good at asking questions. Less than half the people can ask a good question in a meeting." "And then how often do you look at the answer and get really frustrated before it's finished, but you have to pay all this money to watch the seven paragraphs come out...?" "My favorite is the output of it is designed for probabilistic programming." "Instead of saying, 'Is this a customer service question? Then route to customer service, otherwise route to general help desk,' it's, 'This is 80% customer service.' That's exactly simulation." "There's 50 years of computer science research in literally probabilistic if statements. Suddenly the coolest place to be in computer science is gonna be probabilistic programming, which was all of computer science in the 1960s and '70s." "This is not all new. It's gonna be very interesting to dust off all of that work, because it's exactly what's going on. An if statement now is if X percent, not if always." "The way that Jev works is it's a custom programming language almost, which is: here's the prompt, come back with a percentage. And then you put that in the if statement." @stevesi
Box CEO Aaron Levie, Steven Sinofsky, and Martin Casado join Erik Torenberg to discuss "We Must Pace the Frontier," the upcoming AI election in 2028, and Jev: They argue that most of today's AI regulation debate is happening before anyone has defined the risks being regulated. Every prior wave, from computer viruses to aviation, built its safety standards after learning how the technology actually failed. Then it gets concrete. Agents don't get tired, run at enormous scale, and probe systems in ways employees never could, which may mean rethinking permissions, authentication, and the security stack itself. They close on why AI innovation may increasingly happen outside the frontier labs, in the software built around the models. 00:50 "We Must Pace the Frontier" 03:50 Do the labs believe their own x-risk talk? 06:49 If it's existential, nationalize it 11:32 "You're asking us to regulate you?" 12:08 2028 as the AI election 18:50 Tech never learned to navigate regulation 25:50 The law that came from one 1983 hack 28:39 Noam Brown's heat exfiltration idea 32:30 Cold War covert channel stories 34:35 Agent swarms look like a DoS attack 41:25 Sinofsky's fear: GDPR for AI 46:20 No jets if the FAA started in 1910 48:38 Jev: decision engines vs. chatbots 53:01 Labs build beings, software needs tools YouTube: piped.video/TLJNJDf2XGo @levie @stevesi @martin_casado @eriktorenberg
24
33
300
77,423
mei retweeted
Introducing jevgrep - a research agent CLI powered by jev from @typesafeai that reduces your coding agent cost by 40% (verified on SWE-bench) Make sure to use the built in skill so your coding agent knows to use jg for context collection github.com/dzhng/jevgrep
99
176
2,555
153,365
mei retweeted
Martin Casado on why the labs missed Jev: they're building beings that speak, and software needed a model that chooses. "LLMs were text in, text out. They generate text, and they came from chat... We've spent the last few years trying to take this thing that spits out text and cram it into a traditional program... It's just been super janky." "Jev basically said, 'Generating text as output is very expensive, but it's also more complicated than you need... If you give us a set of options, we'll choose the best option. We can do that incredibly fast, incredibly cheaply, but also with much more accuracy because we can train just for this.'" "This has probably been the fastest adoption of an AI model since ChatGPT. It's been remarkable because we were all primed for this." "[The labs] are trying to create beings, and beings speak. If you're trying to create God, God speaks in natural languages. This is really about something that's for traditional software." @martin_casado
Box CEO Aaron Levie, Steven Sinofsky, and Martin Casado join Erik Torenberg to discuss "We Must Pace the Frontier," the upcoming AI election in 2028, and Jev: They argue that most of today's AI regulation debate is happening before anyone has defined the risks being regulated. Every prior wave, from computer viruses to aviation, built its safety standards after learning how the technology actually failed. Then it gets concrete. Agents don't get tired, run at enormous scale, and probe systems in ways employees never could, which may mean rethinking permissions, authentication, and the security stack itself. They close on why AI innovation may increasingly happen outside the frontier labs, in the software built around the models. 00:50 "We Must Pace the Frontier" 03:50 Do the labs believe their own x-risk talk? 06:49 If it's existential, nationalize it 11:32 "You're asking us to regulate you?" 12:08 2028 as the AI election 18:50 Tech never learned to navigate regulation 25:50 The law that came from one 1983 hack 28:39 Noam Brown's heat exfiltration idea 32:30 Cold War covert channel stories 34:35 Agent swarms look like a DoS attack 41:25 Sinofsky's fear: GDPR for AI 46:20 No jets if the FAA started in 1910 48:38 Jev: decision engines vs. chatbots 53:01 Labs build beings, software needs tools YouTube: piped.video/TLJNJDf2XGo @levie @stevesi @martin_casado @eriktorenberg
23
28
280
69,778
mei retweeted
红杉的一位合伙人 Pat Grady 昨天直接把他们和自己的一个 LP meeting(Boston College)的内容录制发了出来,信息密度非常高,强烈推荐大家去看一下: - JEV 在过去的 7 天里从零增长到了 100 million 的收入。在此之前 The Information 刚刚爆出来,他们现在估值是在 10 billion。 - 现在 AI 公司的增速,比历史上的都要快很多。Instinct 发布以来,保持了 day over day 10% 的增长率。 - 200 到 700,指的是某一个 high-value knowledge work company 从去年年底 200M 增长到今年预测 700M 的收入。(后面那个 2 到 50 和 0 到 70%,他说还是跳过吧,可能还是有些敏感。) - 在 AI 公司,现在有一个连融两轮的行业惯例或风气:把"陪你建公司的"和"只出钱的"拆开。他拿红杉内部的数据(过去 12 个月里的 7 个案例)来举例: 红杉作为陪建的那一方进场,平均的投后估值是 1.1 亿;而仅仅一个月后,只出钱的那一方给的下一轮平均估值就达到了 34 亿。他自己的说法是以前从没见过,这就是泡沫。 - 很多企业想 own 自己的 intelligence:基础模型厂商只给你 Pareto frontier 上有限的几个点(大中小几档),但企业自己的 workload 需要的往往不是这几个点,所以要自己训、自己部署、专门为自己的 workload 优化。这也是为什么他说专用架构的 lab 在跑赢通用的 lab。 - 应用层公司每 4 个月就要 reinvent 一次自己,因为底层模型一直在变,地板一直在抬。 - 组织从 hierarchy、command and control 往 network of agents 走:AI 负责信息流转,人相对自主地工作。他说没有 Jack Dorsey 几个月前那篇帖子说得那么夸张,但方向就是这个。 - labs 内部:模型去年大部分时间已经在自己造自己,最近才意识到 alignment 没被当回事;都在押 custom silicon、新架构和 continual learning,算力全员紧缺;为了抢 API token,一边降价,一边派 consultant 去给 Fortune 500 定制吃 token 的工具 - 模型能力和实际落地之间有很大的 diffusion gap,这就是应用层的机会。 - hyperscaler 今年开始借钱做 capex,不再靠 free cash flow。 而且我特别喜欢他这种说话方式,他每句话就是只说一遍,说过去就过去,所以信息密度很高。推荐大家可以直接去看原视频, 总共视频时长只有 15 分钟:loom.com/share/c016702964a04…
The @BostonCollege Investment Committee (an LP and my beloved alma mater) asked for a few thoughts on what's happening in AI. I recorded a test run yesterday morning and then shared it with my partners, who encouraged me to share it more broadly... so here you go! This is not a sales pitch, it's just a reflection on what we're seeing. And it wasn't intended to be shared, so please pardon the rough edges. loom.com/share/c016702964a04…
44
199
1,737
458,526
mei retweeted
The G in "AGI" stands for Jev
104
29
1,150
41,484
mei retweeted
Jev by @typesafeai went from 0 to $100m in 7 DAYS...WTFFFF
The @BostonCollege Investment Committee (an LP and my beloved alma mater) asked for a few thoughts on what's happening in AI. I recorded a test run yesterday morning and then shared it with my partners, who encouraged me to share it more broadly... so here you go! This is not a sales pitch, it's just a reflection on what we're seeing. And it wasn't intended to be shared, so please pardon the rough edges. loom.com/share/c016702964a04…
10
7
304
76,481
Jev for transaction categorization 😆 Bye bye quickbooks.....
54
79
2,883
141,712
SO MUCH THIS I believe that one of jev's greatest benefits to automation will come from resurrecting architecture best practices: state management, encapsulation, abstraction !!!! and combining them with ML's best practices: measure/evaluate, use calibration/uncertainty (adding screenshot b/c I don't know how to quote 2 posts 🤦)
46
40
530
37,540
Replying to @kieranklaassen
some of our biggest prod successes are in this regime - tune-able RAG
8
5
206
21,698
mei retweeted
DSPy 3.4.0 out, with: (1) native support for Jev and System One models - in the timeless DSPy syntax. (2) a brand new optimizer, ReAnchor, specifically for calibrating outputs with confidence. (3) lightning-fast import speeds for the LLM abstraction, via sibling library LM15)
DSPy 3.4.0 was just released! This release includes native support for Jev and System one models inside of DSPy! Use it with compatible signatures. This release also includes a brand new optimizer, ReAnchor, specifically for calibrating outputs with confidence.
9
33
207
12,262
what’s that new model everyone’s using again?
Lost Internet
1
2
36
6,605
mei retweeted
Jev is 238x cheaper than Astra on input, and output is free. But the price isn't where you lose money. The architecture is. 10k token state x 1,000 decisions: Astra: $100.00 Jev: $0.42 The setup that keeps it that cheap: 1. One call, every question State is billed once. A 13 question briefing batched into one call ran 12.2x cheaper and 10x faster with the same answers. 2. Route in code Ask the speculative questions too, then ignore the answers you don't need. 3. Act on confidence Read only actions run at 0.5. Destructive ones wait for 0.9. 4. Keep math out of the model Counting, dates and arithmetic go to code. Jev only does the judgment. The full setup with the exact config is below.
23
15
142
17,732
Observe the art of the @deel. And they came with receipts 💅 Keep reading to find out how it's done.
19
8
87
51,555
Deel put Jev head-to-head with frontier LLMs across multiple use cases: up to 59× cheaper, up to 4× faster, and quality at parity or better on 6 of 8 checks. TypeSafe opened Jev to us on Sunday. Expected cheaper and faster, got both. The real shock? Accuracy surged. Repeat questions: 70% → 97% Expense categories: 50% → 86% Escalations: same catches, fewer false alarms
Observe the art of the @deel. And they came with receipts 💅 Keep reading to find out how it's done.
23
10
165
30,224
mei retweeted
"JEV-as-a-Judge: Accept When Confident, Escalate When Unsure" This paper shows you can just use JEV for every evaluation instead of expensive LLM. JEV basically acts as a cheap first-pass judge, returning both a verdict and how confident it is. When confidence is high, keep the answer. When it’s low, escalate to a stronger LLM. This simple routing keeps ~99% of GPT-6’s accuracy while reducing evaluation cost by a lot. alphaxiv.org/abs/2609.26550
57
229
1,760
247,044
mei retweeted
Jev just became the fastest-adopted model in AI history. here's what people already built with it 1. jev-ultrafast - the browser agent that picks every click itself, only calling a text model when it actually needs to type something. found real flight results in 7 seconds for $0.0039. 16,758 stars github.com/browser-use/jev-u… 2. jev-trader - real trading bot placing live limit orders on Monad every 300ms block, judged by Jev alone. 1,911 stars github.com/jarrodwatts/jev-t… 3. jev-usecases - production security-operations harness where Jev triages incidents and gates every escalation behind a confidence cutoff before anything touches real infrastructure. zero false escalations in the committed test set github.com/kenhuangus/jev-us… 4. tax-doc-classifier - sorts real IRS tax forms with 100% strict accuracy across 261 forms, at roughly $0.001 a page github.com/kyotofin/tax-doc-… 5. jev-drone - a simulated quadrotor clears a five-station obstacle course by camera alone, Jev judging the situation twice a second github.com/RomanSlack/jev-dr… 6. killmyidea - describe your startup idea, Jev scores it from every angle, then hands back kill, fix, or ship in seconds, not days github.com/monteduro/killmyi… 7. jev-curate - streams Parquet and JSONL rows through typed judgments at 1,500+ rows a second, keeping only what clears the bar github.com/AkashPriyadarshii… 8. pg-jev - a PostgreSQL extension that lets you ask your own database tables plain-English questions and get a real answer back, no SQL required github.com/realZachi/pg-jev eight repos. zero generated words. every single one returns a typed answer against a question someone already defined full setup below, then run the three-question test from the article before you build a ninth
16
35
399
65,826
mei retweeted
Jev is the "Internet" moment for the AI industry It tells your agents and LLMs what to do next, in milliseconds and at almost zero cost If you set it up correctly, you will have the AI engineer’s stack for 2028 In this document, I show you how Bookmark it, then check out the full Jev guide below
7
61
351
51,850