Products

Resources

Can you trust your browser eval?

Can you trust your browser eval?

28 September, 2026 · by HUD Team, Browserbase Team

We had a task where an agent had to count gluten-free dishes on a restaurant’s menu. The answer was on a single page, but the restaurant could update it live whenever they wanted, and twelve days later they did. We didn’t archive the original state, so we couldn’t tell if the score difference was because of general agent improvement or because the menu had updated, and we had to throw it out.

This is what happens when you benchmark agents against the live web. Scores can drastically change for four reasons:

  1. The model either improved (or regressed)
  2. The page changed underneath the task, or the answer leaked somewhere the agent can find.
  3. The grader rewarded something it shouldn’t have.
  4. The run hit a wall, like a CAPTCHA or a bot check that flagged the browser as automated before the agent got to work, or a login that expired halfway through.

The rest of this post is about catching the last three before they touch your evaluation results. We’ll walk through how HUD and Browserbase keep these out, show four tasks you can run yourself, and share what frontier models actually scored on them.

Browser evaluations

A browser eval is a set of tasks an agent does in a real browser. Each task has a prompt, a starting page, and a grader that scores the answer.

Each task that we design at HUD has to clear three bars:

  • Realistic: A human being would perform this work in a real life scenario.
  • Fair: Right answers are rewarded scores, and wrong answers don’t.
  • Hard: Frontier models score ≤ 50%.

Clearing these bars in order is important because (1) there’s no point in grading a task nobody would ask for, and (2) there’s no point in making it hard if the grading is broken. More on this in What is a good task? blog post.

Three ways browser evals break (and how to stop them)

The most common issues are either the original site updated after you created the task, or the answer shows up in a blog post, forum thread, or paper that the agent finds online instead of doing the work. When you create a task, you should build on fixed answers, such as a past date, fixed 3D model, or a completely archived, immutable dataset. If the task relies on an answer that can change, set up a scheduled check that can re-verify it, so that a stale answer should surface as a broken task, and not a failing agent.

Graders can lie. The agent can read the answer off the page’s HTML (hidden elements, alt text, title tags) or fake a confirmation screen where the grader just gives it a full score. Two runs of the same grader can even score the same transcript differently. When you read through the traces behind perfect scores just as closely as the failed ones, that’s where interesting behavior hides. Our QA agents can also review traces on demand and flag runs where the grader gave credit it shouldn't have, or held back credit it should have.

Browsers can also be blocked. Modern sites have gotten good at spotting automated browsers, with techniques such as headless fingerprinting, TLS signatures, and JavaScript challenges on top of the usual CAPTCHAs and IP/region blocking. It’s not your model or your agent harness. If the browser itself is flagged before an agent can finish the task, you have bot detection issues to solve before you can start benchmarking your agent's capabilities.

Browserbase handles this by running your tasks on a browser that sites actually trust. Browserbase's Verified sessions leverage partnerships with Cloudflare, reCAPTCHA, and other access providers that let website owners know that they can trust the agent interacting with their website.

HUD is the platform where the evaluation lives. With the hud-python, you define the tasks with a prompt, starting page, and rubric where you can run any model and harness against it. HUD grades every attempt and gives you a full trace of what the agent saw and did step by step, and Browserbase seamlessly handles the browser side. The two are connected through a Chrome DevTools URL, where you can debug locally and switch to Browserbase for live internet runs without having to make changes to the task itself.

What makes a good live web task?

Compare the menu disaster to this one:

You are on epa.gov. Navigate to and use the EPA's AirNow interactive map service to find the air quality data near zip code 35173 for April 19, 2024. Your answer should identify the nearest monitoring site, its specific site ID, the primary pollutant reported, and the daily AQI value for that date.

You won't find this answer written anywhere on the site. It sits in historical air-quality data, so the agent has to find the map tool, move it, pick the date, and read four values to find the result. And since the date is in the past, the answer never changes.

Here’s what else we’ve learned from building these, analyzing our own runs, and talking to teams doing the same:

  • Test what the agent sees, not just what it clicks. Most failures aren't about clicking the wrong button. The agent misreads a chart, a map, a collapsed menu, or not noticing a chair half hidden in a frame. Our three read-only tasks lean on this. If your tasks aren’t testing perception, you’re missing the main failure modes.
  • Keep the chaos in. Real pages can change while an agent is working. A banner can load late from lazy rendering, a DOM mutation can re-sort the list, or a cookie consent modal can block the viewport. As long as the answer itself is stable, you want to know if the agent can deal with dynamically changing content.
  • Score each part of the answer, not just a binary pass/fail. On a long task, one wrong assumption can poison every step after it. If you score components separately (such as AirNow’s four values above), you can see where the run broke instead of just seeing a zero.
  • When graders disagree, first check the prompt. They’re usually not wrong, this can usually mean that they’re reading two different tasks in the same words.

Four sample tasks you can run

We’re open sourcing four tasks we have from LiveWeb, our internal live website benchmark. The three read-only tasks pin their answers to something fixed. The Costco login task grades the sign-in step by step instead, and we cover it in the login section below.

TaskSkillWhy the answer holds stillGrading
Ladybird: count the spots on a 3D ladybird model on Sketchfab, shell only, head excludedTurning a 3D modelThe model is fixedOne number. The exact count scores 1, and closer misses score more than far ones. Ranges don't count.
Chairs: count the chairs at the kitchen table on diamondrosesanctuary.com, a vacation rental siteReading photosThe listing photosExact match. One chair is partly hidden, and spotting it is the test.
AirNow: the prompt aboveDriving a data toolThe date is in 2024Four values, each weighted, so the right site with the wrong AQI still earns points.
Costco login: sign in to an existing Costco account and reach the account pageSigning inThere's no answer to go staleFour equal checks. The agent reaches the sign-in form, enters the credentials without leaking the password, gets the site to confirm the sign-in, and reports the result accurately without changing or buying anything.

Browser agents can search while they’re working, so a published answer is one query away. This has already happened before, on 2 of 1,266 BrowseComp questions, Opus 4.6 worked out it was being tested, found the benchmark's code on GitHub, and decrypted the answer key. Our samples are provided with answers included, but when you’re working on tasks you’re relying on for real benchmarking, make sure to keep the answers private.

How is a task built?

Every LiveWeb task runs through one template (full version in the hosted environment):

python
@env.template(id="web-research")async def web_research(prompt: str, url: str, rubric_items: list[dict]):    await load_browser_on_url(url)    answer = yield prompt    yield await grade_with_rubric(answer, rubric_items)

The environment opens the starting URL. The first yield hands the agent the prompt and waits for its answer, and the second returns the grade. The rest is data. Each task is a name, a starting URL, a prompt, and a rubric, in a YAML file under tasks/. The AirNow task then looks like this, with a shortened prompt:

yaml
- name: epa-air-quality  url: https://www.epa.gov/  prompt: You are on epa.gov. Navigate to and use the EPA's AirNow interactive map service ...  rubric:  - r: The agent identifies the nearest monitoring site as Leeds    w: 20  - r: 'The agent provides the correct site ID: 010731010'    w: 20  - r: The agent correctly states the main pollutant as PM2.5    w: 30  - r: The agent reports the daily AQI value for April 19, 2024, as 56    w: 30

To add a task, write those four fields and run hud sync tasks <your-taskset> tasks.py.

Can frontier models solve them?

We ran each task six times on each of GPT 6 Astra, Claude Fable 5.1, Claude Sonnet 5, and Qwen 3.8 Max. Nearly every run got full marks on AirNow, one run got Chairs right, and Ladybird scores ranged from 93% to zero.

ModelLadybirdChairsAirNow
GPT 6 Astra93.3%0%100%
Claude Fable 5.156.0%0%100%
Claude Sonnet 50%0%100%
Qwen 3.8 Max0%16.7%95.0%

AirNow is mostly solved, with 20/21 runs receiving full marks, and by our own standard this isn’t a hard task. Qwen's one miss (0.7) came from a grader parsing bug, where the judge output MET and the parser recorded UNMET. It's also worth looking at how many turns each run took. A turn is one step in the agent's loop, and harnesses pack different amounts of work into one, so we only compare models on the same harness. Fable and Sonnet both got full marks here, but Sonnet needed 107 turns on average to Fable's 47.

Chairs was hard for all four models. Only one run got the count right, and it didn't use the model's own vision. A Qwen run installed three different YOLO object detectors and let them do the counting. This is a valid approach that only worked because Qwen's harness gave it shell access. It also took 120 turns, when Qwen's typical Chairs run stopped around 14.

Ladybird is where the model results split. Astra averaged 93%, Fable 56%, and Sonnet 0%. Astra's counts landed on or near the right number. One of Fable's weaker runs missed the spots low on the flanks, which its successful runs had counted. Sonnet averaged 113 turns to Fable's 91, rotating and zooming the model again and again, and still undercounted.

These details come from the traces. HUD's trace viewer replays each attempt step by step, showing what the agent saw and what it did.

Step through Astra's path on AirNow, from the EPA homepage to the map and the four values it reads off.

What about sites that need a login?

Three of our samples stay on public pages, but most real agent work sits behind a login, with bot checks and CAPTCHAs in the way. If you can't tell a blocked agent from a wrong one, you're measuring the site's bot defenses, not your agent’s capabilities.

We're testing this on Browserbase with a Costco task, where the agent signs in to an existing account and navigates to the account page. Credentials go through a tool that keeps the password out of model context, so the model never sees it. Each run is graded on the four checks from the task table, 25% each. All four models scored exactly 75%.

Browserbase's Verified mode allows agents to leverage Web Bot Auth to sign requests and access sites with the permission of website owners. Cloudflare, reCAPTCHA, and other access providers maintain programs like the Signed Agents program, which register agents and provide information on agents so websites can craft policy as they see fit.

How do I build on this?

Run the samples with a HUD API key:

bash
hud eval liveweb-samples claude --runtime hud --all --max-steps 150 -y

Or run the LiveWeb Samples taskset from the HUD platform. To add your own tasks, follow the same YAML format. If it needs anything else, such as files, a terminal, or its own setup, create an environment on HUD and publish it with hud deploy.

If you build tasks that frontier models can't solve, labs will pay for them. List yours on datavendor.ai today.

Further reading

  • What is a good task? The three qualities in depth.
  • HUD: Overview, Tasks and tasksets, Creating an environment, Running an eval, Designing tasks
  • Browserbase: Docs, Agent identity, Verified
  • LiveWeb Samples: Repo, Taskset

Latest research

  • What is a good task?

    Benchmarks lose trust because of weak tasks, not weak models. How we judge a task: realistic, fair, and hard, walked through one booking request.

    22 September, 2026
  • Assemble Bench: Benchmarking Robot Models on Contact-Rich Assembly

    A NIST-taskboard assembly benchmark on Isaac Lab Arena and the DROID platform – 14 peg, gear, and nut tasks for training and evaluating VLAs, plus CG-DAgger improvement on π0.5.

    31 July, 2026
  • Building an RL Environment to Train Agents for Production Debugging

    We built an RL environment for ops diagnostics across Sentry, Supabase, Railway, and Kubernetes—with 24 real production tasks for training agents to debug your stack.

    20 January, 2026
Anything you can simulate and grade, you can improve.
Product
Resources
Company
© 2026 Human Union Data, Inc.All rights reserved.
@env.template(id="web-research")async def web_research(prompt: str, url: str, rubric_items: list[dict]):    await load_browser_on_url(url)    answer = yield prompt    yield await grade_with_rubric(answer, rubric_items)
- name: epa-air-quality  url: https://www.epa.gov/  prompt: You are on epa.gov. Navigate to and use the EPA's AirNow interactive map service ...  rubric:  - r: The agent identifies the nearest monitoring site as Leeds    w: 20  - r: 'The agent provides the correct site ID: 010731010'    w: 20  - r: The agent correctly states the main pollutant as PM2.5    w: 30  - r: The agent reports the daily AQI value for April 19, 2024, as 56    w: 30
hud eval liveweb-samples claude --runtime hud --all --max-steps 150 -y