the SDK for browser agents, built by @browserbase join the discord: discord.gg/stagehand

chrome devtools
Filter
Exclude
Time range
-
Minimum likes
Stagehand retweeted
How do you actually build an effective harness with Claude? We had Thariq (@trq212) from Anthropic at Navigate 2026 to talk about "Unhobbling Claude", the difficulties and processes on how to build agents and harnesses.
7
5
50
10,487
GPT-6 Luna is a step function improvement from it's predecessor scoring 11% higher in accuracy on our browser agent tasks. GPT-6 Sol scores 2% better in accuracy, but costs 20x more per task than Luna.
GPT-6 Sol and Luna just landed in Astra’s orbit. Both launch today with API prices 50% lower than GPT-5.6. Build with Sol. Scale with Luna. To production and beyond.
5
2
15
1,449
Stagehand retweeted
Stripe's new WebMCP endpoint lets your agents make full purchases on the web. Using @Stagehanddev you can navigate to any Stripe checkout page and use their new WebMCP tools to fill the checkout and fetch order details.
Stagehand can use Stripe's new WebMCP tools in checkout pages. Run page.tools() to see what WebMCP connectors a page exposes. With Stripe, you can use Stagehand to add items to cart and complete the checkout with native tools.
3
5
42
5,448
Stagehand can use Stripe's new WebMCP tools in checkout pages. Run page.tools() to see what WebMCP connectors a page exposes. With Stripe, you can use Stagehand to add items to cart and complete the checkout with native tools.
Hi internet! We went ahead and upgraded all @stripe hosted checkout pages for 7.8M businesses—0.45% of the world’s GDP—with WebMCP so agents can more efficiently buy. Our evals show this reduces checkout latency by 39% and token usage by 42%. Learn more: stripe.dev/blog/how-stripe-i…
3
4
28
6,599
During early testing on our Browser Agent Evals, Opus 5.5 completely crushed it's predecessor Opus 5, but also Fable 5.1 & 5 in accuracy, speed, and, cost. Compared to Opus 5 in the Claude Code harness it's roughly 5x cheaper per task, and 2x faster.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
8
2
18
1,350
Everyone can build Agentic Commerce using the @link and Browse CLI powered by Stagehand!
it seems like agentic commerce is finally solved with stripe link. and we're excited to partner with stripe to enable all agents to transact online. your agent can now natively make payments on the web using link cli + browserbase. one shot prompt: docs.browserbase.com/integra…
3
20
3,006
Muse Spark 1.3 from @Meta tops the charts in our Browser Agent benchmarks using the @mastra agent harness. In this harness, Muse outperforms Opus 5 in cost, speed, and accuracy.
19
7
80
8,040
Introducing our Browser Agent Evals. We benchmark frontier and open models using various harnesses on computer-use tasks. Our evals are fully open source and reproducible via our Evals CLI.
10
11
35
6,665
We got early access to Fable 5.1 and evaluated it extensively across our internal benchmarks. It outperformed all other frontier models (including Fable 5) on accuracy scoring 92.11%, while being 60% cheaper than Fable 5 per task.
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work.
5
5
25
2,765
FYI: Stagehand supports WebMCP as a first-class primitive. just call 𝚙𝚊𝚐𝚎.𝚝𝚘𝚘𝚕𝚜() in your agent! happy hacking :)
justine
Create agent-ready web apps for the WebMCP Challenge → goo.gle/4qpDQYV We're excited to see ChatGPT supporting WebMCP, the experimental standard that lets users and AI agents navigate websites together. In the @OpenAIDevs WebMCP Challenge you can experiment with it and win prizes. Plus, Chrome's @sarah_edo is on the judging panel! What will you create?
2
11
87
12,546
Stagehand caches act(), observe(), and extract() results server-side to reduce LLM costs. With v4, you can configure thresholds, inspect every hit/miss, and disable caching for specific calls.
1
5
284
Stagehand v4 gives you more control over how your browser actions are cached. With Browserbase caching, your scripts automatically run up to 80% faster.
Introducing Stagehand v4: the SDK for browser agents. Playwright was built for testing, we built Stagehand for your agent: with improved context management, self-healing actions, and iframe support.
2
4
35
3,900
Last week we launched Stagehand v4, the fastest and most token efficient way to control a browser. We benchmarked it extensively and found that Stagehand repeatedly outperforms Playwright in every action and in every region on Browserbase.
Introducing Stagehand v4: the SDK for browser agents. Playwright was built for testing, we built Stagehand for your agent: with improved context management, self-healing actions, and iframe support.
4
1
17
1,444