Prerequisites
Set your API token and base URL. Both integrations require at least one proxy URL; unlike a direct API request, the account’s default pool is not used:- scrapy-cloud-browser
- scrapy-playwright-cloud-browser
Run the example
From the project directory:results.json should contain an item with "title": "Example Domain".
Click, type, and scroll
Withscrapy-playwright-cloud-browser, pass an async callable to PageMethod to send page-level CDP commands before Scrapy receives the response:
"playwright_page_methods": [PageMethod(fill_email)] in its request metadata. Replace the selector with an input on that page.
The scrapy-cloud-browser handler does not expose a Playwright page. Use the Playwright integration for these actions.
Use the Surfsky SDK
To manage a session directly from a spider, use the SDK with the Scraping API. This standalone example uses Scrapy to parse the returned HTML. It requires Python 3.12+ and Scrapy 2.13+. Use a separate environment from the Playwright cloud browser integration, which pins Scrapy 2.12:sdk_spider.py outside your configured Scrapy project. This example does not use the cloud browser extensions or download handlers:
proxy to client.session(), and a persistent profile UUID reuses saved state.
Pool settings
A browser serves several requests before it is recycled, so do not assume each response comes from a fresh session. If startup repeats without producing responses, check the API host, token, and proxy URL, then limits. Start with one browser until the spider works.