The agent that checks itself before it clicks

Harsh Chhajer
3m read

You let an agent click through your site to check a flow. It clicks the wrong button once, submits a form it shouldn't have, and now you're explaining to someone why a test run sent a real email to a real customer. The fix wasn't slowing the agent down. It was giving it a second process whose only job is catching the first one.

Anthropic shipped that as the actual product

On August 26, 2026, Claude in Chrome went generally available across Anthropic's paid plans, letting Claude click, type, and navigate through a real browser session using your existing logins. Acting without asking permission at every step is the headline. The interesting mechanism sits right before anything that matters: a separate check reviews the action against what you originally asked for and blocks it if it doesn't match, before a form gets submitted or a file gets downloaded.

That is not a permission prompt. It's a second, different process reviewing the first one's work, running automatically, before the consequence happens instead of after.

Why one process checking itself doesn't work

The reason this needs to be a separate check, not the same agent double-checking its own plan, is the same reason a designer reviewing their own file misses the thing a second pair of eyes catches immediately. A process that generated an action is already committed to that action being reasonable. Asking it to also be the judge of whether it's reasonable means the judge and the defendant share the same blind spots.

Anthropic frames the security case for this in similarly stark terms: the real threat isn't the agent misunderstanding your request, it's content on a page trying to hijack the agent into doing something else entirely. Their own testing reports that pairing content-scanning probes with a separate action-verification classifier brought the measured success rate of these attacks down to 0%. That's Anthropic's own number, from Anthropic's own test, and it should be read as exactly that: a strong internal result, not an independently audited guarantee.

The design version of the same problem

An agent that generates a UI and then reviews its own output has the identical structural weakness. It can't see past whatever pattern it already committed to, because the reviewing pass and the generating pass are the same process wearing a different prompt.

If you're using an agent to build and then grade its own interface, that self-review step is worth roughly nothing as a quality gate. It's the equivalent of skipping the separate check Anthropic just decided a browser agent can't safely operate without.

The one thing to change this week

Whatever agent is producing your design or code output, add one deliberately separate pass before anything ships: a different session, a different prompt, explicitly told to find problems rather than confirm the first pass looks fine. It costs one extra prompt and a few minutes. What it buys is the one property a single process reviewing itself structurally cannot provide.