Skip to main content
Surfsky’s Scrapy integrations use a pool of cloud browsers to fetch pages. The download handler returns an ordinary Scrapy response for your spider to parse. To have a coding agent set this up in your project, give it the setup prompt.
Use human-emulation commands for clicks, typing, and scrolling. Keep Scrapy selectors and Playwright’s navigation and waiting for reading and preparing the page.

Prerequisites

Set your API token and base URL. Both integrations require at least one proxy URL; unlike a direct API request, the account’s default pool is not used:
Create a Scrapy project if you do not have one:

Installation

Configuration

Add to surfsky_spider/settings.py:

Example spider

Save as surfsky_spider/spiders/example.py:
This integration does not expose BROWSER_SETTINGS or FINGERPRINT through CLOUD_BROWSER. Use the Playwright integration below if you need those settings.

Run the example

From the project directory:
results.json should contain an item with "title": "Example Domain".

Click, type, and scroll

With scrapy-playwright-cloud-browser, pass an async callable to PageMethod to send page-level CDP commands before Scrapy receives the response:
Define the function above the spider class, replace the request URL with your form page, and set "playwright_page_methods": [PageMethod(fill_email)] in its request metadata. Replace the selector with an input on that page. The scrapy-cloud-browser handler does not expose a Playwright page. Use the Playwright integration for these actions.

Use the Surfsky SDK

To manage a session directly from a spider, use the SDK with the Scraping API. This standalone example uses Scrapy to parse the returned HTML. It requires Python 3.12+ and Scrapy 2.13+. Use a separate environment from the Playwright cloud browser integration, which pins Scrapy 2.12:
Save as sdk_spider.py outside your configured Scrapy project. This example does not use the cloud browser extensions or download handlers:
Run it:
The SDK stops the browser when the session scope exits, including on errors. It uses your account’s default proxy pool unless you pass proxy to client.session(), and a persistent profile UUID reuses saved state.

Pool settings

A browser serves several requests before it is recycled, so do not assume each response comes from a fresh session. If startup repeats without producing responses, check the API host, token, and proxy URL, then limits. Start with one browser until the spider works.

Stop the sessions

The pool closes its browsers when the crawl finishes. After an interrupted crawl, list active sessions and stop any browser left running, or stop them all:
Otherwise each browser stops after its inactivity timeout, 30 seconds by default.