-
Install Scrapy and create a project
Install Scrapy with pip and generate a project. The examples below use a project named
proxydemo.Shell python -m pip install scrapy scrapy startproject proxydemo cd proxydemo -
Add a proxy middleware
Add this class to
proxydemo/middlewares.py. It sets the gateway on every request that does not carry a proxy yet, and Scrapy turns the credentials in the URL into a Proxy-Authorization header.Python # proxydemo/middlewares.py class ProxonymProxyMiddleware: """Send requests through the Proxonym gateway unless they set their own proxy.""" PROXY = "http://USERNAME:[email protected]:8000" def process_request(self, request, spider=None): request.meta.setdefault("proxy", self.PROXY) -
Enable it in settings.py
Register the middleware with a priority below 750 so it runs before the built-in HttpProxyMiddleware, and set retries and a timeout that suit a proxy gateway.
Python # proxydemo/settings.py DOWNLOADER_MIDDLEWARES = { "proxydemo.middlewares.ProxonymProxyMiddleware": 350, } RETRY_TIMES = 3 RETRY_HTTP_CODES = [429, 500, 502, 503, 504] DOWNLOAD_TIMEOUT = 60 -
Check the exit IP
Save this spider as
proxydemo/spiders/ip.pyand runscrapy crawl ip -O ip.json. The output file should contain an IP address that is not your own.Python # proxydemo/spiders/ip.py import scrapy class IpSpider(scrapy.Spider): name = "ip" start_urls = ["https://api.ipify.org?format=json"] def parse(self, response): yield {"ip": response.json()["ip"]} -
Set sessions and locations per request
A request that sets
meta["proxy"]skips the middleware default. Give a crawl path its own session ID to keep one IP for up to 120 minutes, and add a country, state or city to localize results.Python # proxydemo/spiders/geo.py import secrets import scrapy def proxy_for(country=None, city=None, session=None, lifetime=30): user = "USERNAME" if country: user += f"-country-{country}" if city: user += f"-city-{city}" if session: user += f"-session-{session}-lifetime-{lifetime}" return f"http://{user}:[email protected]:8000" class GeoSpider(scrapy.Spider): name = "geo" start_urls = ["https://ipinfo.io/json"] def parse(self, response): yield {"rotating": response.json()} # Three follow-up requests that share one sticky IP in Paris session = secrets.token_hex(4) for n in range(3): yield scrapy.Request( "https://ipinfo.io/json", callback=self.parse_sticky, dont_filter=True, meta={"proxy": proxy_for(country="fr", city="paris", session=session)}, cb_kwargs={"n": n}, ) def parse_sticky(self, response, n): yield {"n": n, "sticky": response.json()} -
Tune retries and concurrency
Scrapy retries timeouts, refused tunnels and the status codes in
RETRY_HTTP_CODES. A refused tunnel is logged with the gateway status, so fix the cause instead of retrying when you see 407, 402 or 403.Python # proxydemo/settings.py CONCURRENT_REQUESTS = 32 CONCURRENT_REQUESTS_PER_DOMAIN = 8 AUTOTHROTTLE_ENABLED = True AUTOTHROTTLE_TARGET_CONCURRENCY = 4.0
Frequently asked questions
Does Scrapy support SOCKS5 proxies?
Not natively. Scrapy talks to HTTP proxies, which is all you need for Proxonym: the gateway on port 8000 tunnels HTTPS through CONNECT and accepts every targeting and session parameter. Keep the SOCKS5 endpoint on port 1080 for tools that require it, such as non-HTTP clients, and point your spiders at the HTTP endpoint.
Why does Scrapy log a CONNECT tunnel error?
The gateway refused to open a tunnel, and the log line includes its status code. A 407 means wrong credentials, 402 an empty balance or an expired plan, and 403 a target blocked under the acceptable use policy. A 502 or 504 is a transient upstream failure that the retry middleware handles for you.
How many concurrent requests can a spider send?
Proxonym does not cap concurrent sessions, and each account has a burst limit of around 10,000 requests per second. The practical limit is usually the target site, so start with 16 to 32 concurrent requests, enable AutoThrottle, and raise concurrency gradually while you watch error rates and response times.
Do I need a proxy rotation plugin for Scrapy?
Not for the residential gateway. One endpoint rotates IPs on the network side, so there is no list to cycle through. Rotation plugins are designed for static proxy lists, such as a shared datacenter IP list, where the spider itself has to pick a different address from the list for each request.
