Entrepreneur. Advisor. AI and cybersecurity. Current: CEO/founder @polarskyinc. Prev: ceo/founder @ambassadorlabs; product @duosec; corp dev/strategy @rapid7.

Boston, MA
Pinned Tweet
Companies are giving AI assistants access to internal data, SaaS apps, and tools faster than they’re building the security model around that access. The question isn’t just “is the model secure?” It’s: What can the assistant access? What can it do? On whose authority? And what happens when it gets something wrong? I spent the past year talking to CISOs, security practitioners, and AI experts about how they’re approaching this. I couldn’t find a practical framework for securing ChatGPT, Claude, Copilot, and other AI assistants in the enterprise, so I wrote one. Preflight starts with the threat model, then maps the controls you should deploy and the order to deploy them. Read it: polarsky.ai/guide/
4
73
Your biggest insider risk isn't human any more. It's the AI assistant that can access what your employees can access. As enterprises connect AI assistants to internal systems, those assistants inherit access to documents, customer data, workflows, and sensitive business context. The result is a new class of privileged insider: non-human, highly connected, and capable of operating at machine speed. And they don't need to break in. They already have legitimate access. Security architectures weren't designed for this. Our new guide introduces a threat model and controls framework for these connected AI assistants. polarsky.ai/guide
25
Re: #AI lights-on/lights-off: TBH, I think the interesting question is what protocols do you have in place to ensure quality. Humans are only one protocol, and framing it is lights-on/lights-off ignores all the other ways we can (and should!) manage quality: static checks, AI spec reviews, human spec reviews, ...
1
33
Fable 5.1 is _much_ easier to understand compared to the other Anthropic models. Thank goodness!
1
51
Claude UI is clearly targeted towards technical users: there are explicit widgets for deterministic control (e.g., scheduling). ChatGPT is targeted towards more general users: no matter what you do, the generic interface is one box that you type into.
2
75
And @tryramp just launched router.com. My take: inference routing is migrating from "developer infrastructure" to "enterprise procurement". AI inference is now a financial primitive. Very curious to see how this evolves.
OpenRouter is joining Stripe. We started OpenRouter with a simple mission: intelligence should be multi-model. Today, we are the largest AI marketplace & gateway, processing 10T+ tokens daily on 400+ models. Joining @Stripe gives us the opportunity to accelerate that mission.
2
94
Inspired by deepmind.google/blog/weather…, I wanted to see my 15 day weather forecast. Claude couldn't find anything, so I had it code up github.com/richarddli/weathe… -- gives a 15 day weather forecast based on WeatherNext 2 for any US zip code, and then confidence cross-checked with ECMWF.
4
63
Richard Li retweeted
At the @aiDotEngineer World Fair, I sat down and dumped my brain on a podcast. Here are my latest ponderoos on these topics… ✨ Why code doesn't need to be readable by a human anymore; it needs to be explainable to one. ✨Why frontier intelligence isn't required for most tasks, and why some of the sharpest engineers I know run 20 concurrent $300/year subscriptions instead of one frontier plan. ✨ Why I haven't hand-written code in over two years, and why Git is perhaps already end of life. ✨ Why porting between programming languages is now nearly free, and what that means for which languages survive.
8
12
106
32,126
I'm really interested if someone can do a _writing_ eval and not just an "intelligence" or "coding" eval. Subjectively Sol seems to be a better writer than Fable 5. Haven't used Opus 5 enough yet to tell.
Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task We supported @AnthropicAI to evaluate Claude Opus 5 ahead of release: it sets the highest GDPval-AA v2 and AA-Briefcase scores so far. Opus 5 (max) scores 61 on the Artificial Analysis Intelligence Index, effectively tied with Claude Fable 5 (max, 60), and ahead of GPT-5.6 Sol (max, 59), Kimi K3 (57), and Claude Opus 4.8 (max, 56) Key takeaways: ➤ New leader in agentic knowledge work: Claude Opus 5 (max) scores 1861 Elo on GDPval-AA v2, >100 points ahead of Claude Fable 5 and GPT-5.6 Sol (max). On AA-Briefcase, our proprietary agentic knowledge work benchmark, it scores 1720 Elo, +146 ahead of Fable 5. These benchmarks test the ability of models to produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup ➤ Joint first place on the Coding Agent Index: Claude Opus 5 (xhigh) with Claude Code leads the Artificial Analysis Coding Index, including the highest score on SWE-Atlas-QnA ➤ Frontier intelligence with reduced cost: Claude Opus 5 (max) costs $2.03 on average per Intelligence Index task, below Claude Fable 5 (with fallback) at $2.75, but still above Claude Opus 4.8 (max) at $1.80 and Claude Sonnet 5 (max) at $1.53. However, at high and xhigh reasoning efforts Opus 5 can outperform both Opus 4.8 and Claude Sonnet 5 at a lower cost per task ➤ Frontier agentic terminal use: 89% on Terminal-Bench v2.1 at max effort, roughly in line with the leader, GPT-5.6 Sol (xhigh) ➤ Outperformance on scientific reasoning: Along with leading agentic performance, Claude Opus 5 scores 53% on Humanity’s Last Exam in line with Fable 5; on CritPt, a frontier physics evaluation developed by Argonne and UIUC researchers, it also matches Fable 5 but sits behind GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra ➤ Factual knowledge still lags Fable 5: As expected from the models’ size classes, Opus 5 still has lower factual knowledge on AA-Omniscience than Fable 5. It improves +7 points on AA-Omniscience Accuracy over Opus 4.8, but answers more often when uncertain - its hallucination rate rises +14 points to 50% ➤ Improving efficiency, but only on the Intelligence vs. Cost per Task Pareto frontier at high Intelligence levels: Opus 5 outperforms Fable 5 at lower cost, but at lower effort levels it sits just behind the GPT-5.6 family on the Intelligence vs. Cost per Task frontier Other model details: ➤ Context window: 1 million tokens (equivalent to Opus 4.8) ➤ Pricing: As with recent Opus launches, tokens cost $5/$25 per million tokens of input/output; cache pricing remains at a 25% premium for cache writes ($6.25 per million tokens) with 5-minute time to live, and 90% discount for cache hits ($0.50 per million tokens) ➤ Five effort settings (low, medium, high, xhigh, max), and support for server-side fallback as with Fable 5. Intelligence Index evaluations were run with Opus 4.8 fallback enabled
1
173
Richard Li retweeted
We removed ~80% of the Claude Code system prompt for our newest models, this is what we've learned about writing system prompts, skills and Claude.MDs for them.
Article

The new rules of context engineering for Claude 5 models

I’ve written previously about how to best prompt the newest generation of Claude 5 models and work with them iteratively to discover what you want to build. But when you send a message to Claude, the

467
1,904
16,303
4,926,878
gpt-5.6 sol is a _much_ better writer than Fable. Fable can be a genius and do great things on architecture and analysis but sometimes it's incomprehensible (but feeding Fable into GPT-5.6 sol works!)
80
Amazing for @vercel; I'm wondering how @dagster will evolve post-acquisition under @PrefectIO without Pete and Nick.
I’m excited to welcome two legends of developer tools, Pete Hunt (@floydophone) and Nick Schrock (@schrockn), to Vercel. Pete was one of the pioneers of @reactjs at Meta. He made an early bet to power Instagram Web with ⚛️ React, evangelizing it internally and externally. He will be running Frameworks and leading @nextjs. I couldn’t imagine a better person to lead React’s most popular framework to even greater heights. Nick co-invented @graphql, solving some of the gnarliest data infrastructure and access issues at Facebook scale, with a delightful developer experience. He will be working on Agentic Developer Experience, solving the problem of enabling the next billion agents and leading the way to a future of self-improving software. It’s a dream-come-true for a founder of a startup to welcome engineering minds of this caliber who are also wonderful humans. You probably want to work with them, and they’re hiring 😁. Their DMs are open, from job applications to bug reports!
1
247
I've always believed that DevEx is critical to a high-functioning organization, and have learned a lot from thought leaders like @nicolefv @RealGeneKim and @matthewpskelton. Last week, spending time with our #AI team in Los Angeles, I realized that those same principles need a different implementation for scientists. SciEx != DevEx. We had built an extension point around APIs and services, so that scientists could build and rapidly run experiments. That was wrong! We had built a contract around behavior, but what scientists want is a contract around data. My big takeaway: a good platform abstraction should match the user's unit of work, and understanding this is fundamental (and obvious in hindsight). We're rebuilding our SciEx now, and am anxious to see how this works in the real-world.
2
90
Richard Li retweeted
We're extending Claude Fable 5 access on all paid plans, as well as keeping Claude Code’s weekly rate limits 50% higher, through July 19.
6,802
7,143
73,283
28,201,269
GPT-5.6 Sol and Luna are ahead of Terra at every point on the Intelligence vs Cost per Task chart. GPT-5.6 Luna stands out as a particularly cost efficient model Charting the Artificial Analysis Intelligence Index shows the trade-off between intelligence and Cost per Intelligence Index Task. Across reasoning efforts, each GPT-5.6 model pushes past GPT-5.5 on the Pareto frontier (excluding non-reasoning). However, Luna and Sol are always ahead of Terra. This means for any Terra effort level, there is a Luna or Sol effort level that is more intelligent at no extra cost, or as intelligent at lower cost.
122
288
3,189
702,553
I love the @ArtificialAnlys folks and lots of respect to them. One thing though is that this suggests that Fable 5 is on par with GPT-5.5 on index score, and in my experience, this is not the case. The zeitgeist is always hard to read but I'm sure others would agree too.
For agentic coding, one can say: - Unless you need Terra Ultra perf, it's always better to use a Luna model with higher effort setting (same or better performance but cheaper). - Forget everything below Sol High, use Luna with higher effort settings here - Forget Sol Extra High, use Terra Ultra here - The extra cost of Sol Ultra is probably not worth it over Max
2
249
What I'm most curious about is how GPT-5.6 compares to GPT-5.5 on general research & feedback-type tasks. GPT-5.4/5.5 added a lot of verbosity and I noticed a lot of gratuitous advice.
GPT-5.6 Update - GPT-5.6 is not launching today. The rollout has slipped again. - It's still expected to launch this week, with Thursday currently having the highest chances.
1
368
🔥
A quick update on the future of the `transformers` library! In order to provide a source of truth for all models, we are working with the rest of the ecosystem to make the modeling code the standard. A joint effort with vLLM, LlamaCPP, SGLang, Mlx, Qwen, Glm, Unsloth, Axoloth, Deepspeed, IBM, Gemma, Llama, Deepseek, microsoft, nvidia, internLM, Llava, AllenAI, Cohere, TogetherAI.....
1
116
#Agents aren’t the future—they’re already here. But building them? That takes a whole new stack. 🧠 AI ⚙️ Durable execution 🧱 Frameworks 🗂 Context 🛠 Actuators Check out the breakdown (with @jflomenb and @Wing_VC): 🔗 wing.vc/content/the-agentic-…
1
1
4
1,428
This is the agentic runtime stack—what’s needed to run autonomous, goal-directed agents in production. #AgenticAI #LLM #AIinfrastructure
58