94 points wittydeveloper 19 hours ago 24 comments

wittydeveloper 19 hours ago | parent

We built Stagehand 2 years ago (24k stars and 4M monthly npm downloads) and recently fixed its biggest flaw: round-trip latency.

Every action performed requires a round trip between your script and the browser (short when running locally but increased when running in the cloud). We also saw multiple posts complaining about the eager token appetite of Playwright MCP.

For this reason, we rebuilt Stagehand from the ground up and shipped v4, where Stagehand controls the browser from an extension automatically loaded upon your browser startup.

Stagehand v4 comes with batch command support, dedicated token-efficient methods `act()` and `extract()`, and a brand new architecture making it 2x faster than Playwright and 80% more token efficient.

You can see for yourself by looking at our benchmarks, comparing its performance across a dozen models (frontier and open weights) and tools (Codex, Claude Code, and more): https://www.stagehand.dev/evals

Ask me anything!

cl685 19 hours ago | parent

what did you lose compared to CDP (e.g. cross-origin iframes, downloads running in envs where you can't load extensions)?

wittydeveloper 19 hours ago | parent

It still uses CDP but communicates from an extension within the browser instead of a script running in a separate runtime or, worse, in a separate region.

youngtaff 18 minutes ago | parent

Are you using CDP over web sockets or via a pipe?

cl685 19 hours ago | parent

lowk why not just do astra computer use

smpandya 19 hours ago | parent

I tried this myself - my conclusion was the overhead of parsing a screenshot, generating an action, and being limited to headful mode is much less efficient than reading accessibility trees & generating CDP commands.

basically, browser automation is a closer-to-the-metal abstraction than computer use, allows more flexibility, and ends up being much cheaper at scale!

vishalanton 19 hours ago | parent

Astra with computer use seems to burn a ton of tokens though. Stagehand seems more token efficient.

ChemSpider 24 minutes ago | parent

If agents need to run on the desktop or logged into a user's real browser, the alternative would be Ui.Vision MCP. But stagehand is designed for backend use, so apples and oranges.

alyssamaru 19 hours ago | parent

How much does the harness really matter for evals?

wittydeveloper 19 hours ago | parent

A lot, especially for performance. That's why we built our own benchmarks that account for both the model and the harness. For example, with Claude Opus 5, the accuracy gap can be up to 3% and performance up to 200ms, depending on whether you're using Deep Agents, Eve, or Fx.

More details here: https://www.stagehand.dev/evals

vishalanton 19 hours ago | parent

So for an enterprise with a 1,000 test Playwright suite, does this basically mean ~2x faster CI times? That would be huge.

wittydeveloper 18 hours ago | parent

Exactly, as Stagehand now runs inside the browser, you'll save on the round trip. Also, enabling batch actions will further accelerate your test suite.

pixelstack 6 hours ago | parent

judging from their documentation looks like promised increase only achievable with their main product

dot_louis 19 hours ago | parent

Do I need to pay for Browserbase to use this?

wittydeveloper 18 hours ago | parent

Nope, Stagehand is open-source and works with local browsers by default.

ishankunam 18 hours ago | parent

seems really cool! although, one question i have is why not keep agent() alongside the new primatives? it seems v4 removed agent() entirely rather than offering it with all of act(), observe(), extract().

wittydeveloper 18 hours ago | parent

We removed agent because so many great harnesses are available in the ecosystem. Instead of keeping it, we decided to make Stagehand v4 better integrated with popular harnesses, both at the Coding Agent level (Codex, Claude Code) and frameworks level (Eve, Deep Agents, Mastra, etc)

bensyverson 10 hours ago | parent

If your needs are simpler, I created a tiny headless WebKit browser specifically for agents called Sleepy Hollow [0]

[0]: https://github.com/bensyverson/sleepyhollow

noir_lord 7 hours ago | parent

I love the name.

skybrian 28 minutes ago | parent

What would be the Linux equivalent?

shashanoid 6 hours ago | parent

no matter how much faster you make.. playwright is playwright. Dead bot giveaway.

swingboy 1 hour ago | parent

A lot of folks use Playwright for UI testing.

ulrikrasmussen 6 hours ago | parent

Looks very useful, and I like the caching idea which I think makes it interesting for self-healing CI tests.

How does it determine when a cached act() fails and has to be re-evaluated by the LLM? And in particular, if the cache is saved in the cloud (Browserbase?), won't this lead to a lot of cache churn if used in CI pipelines where different versions of the site are running against the same cache?

Also, is there a technical reason why the cache couldn't just be a local file that's checked in along with the script but must be provided by Browserbase? If it was, devs could heal failing tests locally using LLM calls, while CI runs entirely deterministically.

tengkahwee 6 hours ago | parent

Would you recommend to use this over agent-browser for general agent-based validation work? Any performance benefit?