Cursor + PuppeteerBrowser fetches arbitrary URL

Cursor-built Puppeteer scraper accepts any URL: how to secure browser fetching

Parsing a URL only proves that it has valid syntax. A server-side Puppeteer browser may visit internal addresses, follow redirects, and load images, scripts, frames, or other subresources from additional hosts. Restrict destinations and protocols before navigation, inspect each browser request, and run the browser in a network-isolated environment with resource limits.

This guide covers check 04: Input validation and error handling; check 06: Performance and scalability in the Zenveus Production Readiness Standard.

For builders

What this means, in plain words

A headless browser loads more than the first URL you give it. A public page can redirect it or make it fetch images and frames from private destinations. Isolate the browser’s network and check every request, then prove the internal test service received nothing.

A scoped repair request

Using an AI builder? Paste this

Use this prompt in Lovable, Cursor, Replit, or Claude Code with the relevant server files available.

Audit my Cursor-generated Puppeteer scraper for SSRF across navigation, redirects, frames, images, scripts, and DNS changes. Add per-request validation plus isolated network egress controls, byte limits, timeouts, concurrency limits, and guaranteed browser cleanup. Work on a branch with synthetic data and mocked external services. Show the smallest diff, identify required adapters and deployment settings, and add allowed and denied tests that prove side effects cannot happen before checks pass. Do not disable security checks to make a test pass.

The same failure may appear as

  • Puppeteer accepts any submitted URL
  • Scraper follows a redirect to a private IP
  • A public page loads an internal iframe or image
  • Headless browser consumes excessive time or memory

Find the failure layer

Run these checks before rewriting anything

Each check removes a class of causes. Keep the first failing result, its timestamp, and the production log beside it.

01

Navigation input

Find the request field passed to `page.goto` or equivalent and record checks performed before a browser is launched.

If this failsSyntax validation allows unwanted hosts; restrict schemes, destinations, and resolved addresses before navigation.
02

Secondary requests

Inspect top-level navigation, redirects, frames, scripts, images, and fetch/XHR requests. A safe landing URL can load an unsafe subresource.

If this failsUnchecked secondary requests bypass the first-URL guard; intercept every request and enforce worker egress restrictions.
03

Network fixture

Use a fake internal service and disposable browser container. Verify that direct private addresses, redirects, and pages embedding internal resources are denied before contact. Never test against a live third-party target.

If this failsAny mock internal request means the network boundary failed; trace the redirect or subresource and fix both policy layers.
04

Browser limits

Set timeouts, maximum pages, response bytes, and concurrent sessions. Close the browser even after errors.

If this failsA stuck or leaked browser can exhaust workers; bound execution and close resources in a finally block.

Ranked diagnosis

Common root causes, in the order we would test them

01

URL syntax is mistaken for safety

The server accepts any URL that parses before calling the browser.

02

Redirects and subresources escape validation

Frames, scripts, and images reach hosts not present in the initial URL.

03

Browser shares server network

Headless navigation can access internal addresses if egress is unrestricted.

04

No resource limits

A page can keep a browser busy or consume large responses.

Step-by-step repair

How to fix this in your app

Edit app/api/scrape/route.ts and the browser creation helper. Place the browser worker behind an outbound network policy; request interception is an additional application check.

Production safety ruleNever disable access controls, expose service keys, or add wildcard CORS as a routine shortcut.

01

Choose the destinations the product actually needs

For a fixed-site scraper, allow exact HTTPS hosts and the required port. Reject userinfo and unexpected protocols. If the product accepts arbitrary public sites, use an outbound gateway that validates and constrains the connection destination.

02

Intercept every browser request

Enable interception before navigation. Apply the policy to main requests, redirected requests, frames, scripts, and images. Check isInterceptResolutionHandled before resolving a request when multiple listeners exist.

03

Isolate the browser worker

Deny loopback, private, link-local, metadata, and other internal destinations for IPv4 and IPv6 at network level. Use a fresh browser context, no ambient credentials, and a patched browser with its sandbox enabled.

04

Bound the job and close resources

Set a navigation timeout and worker-level memory, concurrency, response-byte, and total-duration limits. Close the browser in finally. Test direct navigation, redirects, and embedded resources against disposable fixtures.

Implementation example

Puppeteer: exact-host request interception

const allowedHosts = new Set(['docs.example.com']);
await page.setRequestInterception(true);
page.on('request', request => {
  if (request.isInterceptResolutionHandled()) return;
  let allowed = false;
  try {
    const u = new URL(request.url());
    allowed = u.protocol === 'https:' && allowedHosts.has(u.hostname)
      && !u.username && !u.password && (!u.port || u.port === '443');
  } catch {}
  if (allowed) void request.continue();
  else void request.abort('blockedbyclient');
});
try {
  await page.goto(targetUrl, { timeout: 10000, waitUntil: 'domcontentloaded' });
} finally { await browser.close(); }

Replace docs.example.com with the intended host and validate targetUrl before navigation. Host checks cannot prevent DNS rebinding by themselves; enforce connection-level egress restrictions. The snippet assumes an already-created page and browser.

Prove the repair

How to check that the fix worked

Run these checks with synthetic data in your test environment, then repeat the relevant acceptance checks after deployment.

  • Approved fixture: page loads within its limits.
  • Unapproved main URL: rejected before navigation.
  • Redirect, iframe, and image aimed at a disallowed host: blocked; fixture receives zero requests.
  • Timeout or browser error: worker terminates and browser resources close.

If the check still fails

If an internal fixture receives a request, inspect browser service workers, redirects, DNS, and egress policy. Do not claim SSRF is fixed from a top-level URL test alone.

When the built-in AI fix makes it worse

Recover one reproducible failure.

Pause generated changes, restore a known working branch, and capture one failing request with its logs. Change one layer and rerun the allowed and denied checks before proceeding.

Engineering handoff

What Zenveus checks when the quick fix is not enough

We trace one production request through the complete path, isolate the failing boundary, and leave behind evidence your team can repeat.

Browser requests

Top-level navigation, redirects, frames, images, and scripts.

Network isolation

Worker egress policy, resolved destinations, and private-network denial.

Resource cleanup

Timeouts, response limits, concurrency, and browser shutdown on errors.

Fixture logs

Approved-page results and zero contact with the mock internal service.

Typical repair pattern

The scraper cannot contact internal destinations

A disposable browser reaches approved pages while fake internal endpoints see no direct, redirected, or embedded requests.

Evidence left behind
  • Root-cause note
  • Verified production check
  • Rollback and prevention steps

Before the next release

Prevent this failure from returning

Review new browser automation paths as server-side outbound network access.
Keep direct, redirect, and subresource denial tests in the release suite.

Clear answers

Questions teams ask before they touch production

01Does `new URL(input)` prevent SSRF?

No. It validates syntax, not the destination's trust level or what the page later loads.

02If users provide their own AI key, is the scraper safe?

A user-provided AI key changes provider billing exposure. It does not constrain the server-side browser's network access.

03Are top-level URL checks enough?

No. Redirects, iframes, scripts, and images can reach additional destinations.

04How do I know the fix worked in my app?

Run the verification checks on this page against your test environment, then repeat the relevant checks after deployment. Example code needs your app's authentication, data model, and configuration; reading the guide alone does not verify your deployment.

Free next step

Check the boundary before you hand it over

The free tool helps you inspect this symptom. Its result does not establish whether the whole app is production ready.

The Verdict

Know whether the symptom is contained or structural.

We can see the symptom from here. What we cannot tell you from outside is whether it is contained or structural. A scanner collects evidence. A named senior engineer makes the decision. For $299, a named senior engineer reads your code and signs a written Verdict against the nine checks in the Zenveus Production Readiness Standard. The 48-hour clock begins when the required access and context are available. If the report does not give your developer a list they can act on, you do not pay.

The 48-hour clock starts when the required access and context are available.

Scroll to Top