Web Scraping Without Getting Blocked: 12 Practical Techniques
Twelve techniques that keep scrapers running: respect the rules, pace requests, rotate IPs sensibly, stay consistent, back off on errors, cache and monitor success rates.
By the Proxonym team
To scrape the web without getting blocked, make your traffic look and behave like a reasonable visitor: stay within the site's rules, keep request rates modest, spread load across IPs, send complete and consistent headers, keep sessions coherent, back off when the site pushes back, and measure success so you notice problems early. No single trick does it. Blocks come from a pile of small signals, and the twelve techniques below remove them one at a time.
Start with the rules#
1. Respect robots.txt, terms of service and the law#
Read robots.txt and the site's terms before you write any code. Collect public data only, do not scrape behind a login without permission, and minimize personal data, which laws such as the GDPR regulate no matter how it was collected. Ethics and durability point the same way: scrapers that stay polite rarely trigger the aggressive defenses that make scraping expensive. Proxies help you distribute legitimate traffic; they do not make prohibited collection acceptable, and our acceptable use policy applies to every request.
Control your request rate#
2. Rate-limit per domain#
Set a request budget per target domain, not per scraper. Start low, for example one request per second for a mid-sized site, then raise it while you watch error rates. Honor a Crawl-delay directive when robots.txt has one. Rate limits protect the site's capacity, and a site that is not under strain has little reason to block you. Plan schedules around the budget instead of raising it: 100,000 pages at one request per second take about 28 hours on a single domain.
3. Cap concurrency#
Rotating IPs lets you run many requests in parallel, but the target still sees the total. A queue with a per-domain concurrency cap, say 5 to 20 requests in flight depending on the site, prevents bursts that look like an attack. Unlimited concurrency on the proxy side is there to spread work across many domains, not to point all of it at one.
4. Back off on errors, with jitter#
Treat 429 and 503 as instructions to slow down. Honor Retry-After; otherwise wait exponentially longer between attempts, and add random jitter so parallel workers do not retry in lockstep:
import random
import time
import requests
PROXY = "http://USERNAME-country-us:[email protected]:8000"
PROXIES = {"http": PROXY, "https": PROXY}
def polite_get(session, url, tries=5, base=1.0, cap=60.0):
for attempt in range(tries):
try:
r = session.get(url, proxies=PROXIES, timeout=(10, 60))
except requests.RequestException:
r = None # network or tunnel error
if r is not None and r.status_code != 429 and r.status_code < 500:
return r # success, or a 4xx to inspect
if attempt == tries - 1:
break
wait = min(cap, base * 2 ** attempt)
if r is not None and r.headers.get("Retry-After", "").isdigit():
wait = int(r.headers["Retry-After"]) # the server says how long
time.sleep(wait + random.uniform(0, wait / 2))
raise RuntimeError(f"giving up on {url} after {tries} attempts")Look like a consistent visitor#
5. Rotate IPs, with the right pool#
Spreading requests across many IPs keeps each address under per-IP limits. Match the pool to the target: datacenter proxies for lightly protected sites, rotating residential proxies where hosting IPs are filtered. On our residential gateway, a request without a session parameter gets a new IP automatically, so rotation needs no extra code.
6. Send complete, consistent headers#
Default library headers announce a script: a python-requests user agent with no Accept-Language is trivial to filter. Send the full set a real browser sends: User-Agent, Accept, Accept-Language and Accept-Encoding, plus a plausible Referer when you navigate from page to page. Keep the claims consistent, too. A Chrome user agent from a client whose TLS and HTTP/2 fingerprint is clearly not Chrome is a stronger signal than an honest library user agent.
7. Keep sessions coherent#
Accept cookies and send them back, like a browser does. When a flow spans several requests, such as pagination with a cursor, a cart or a login, keep the same IP for the whole flow with a sticky session, then start a new session for the next flow. Rotating vs sticky sessions covers how long to hold an IP.
8. Match geography#
Use exit IPs in the country whose content you want, and make the rest of the request agree: Accept-Language, currency and region parameters, and the timezone in browsers. A German IP asking for English pages priced in dollars is unusual. Where results are local, city targeting such as -country-us-city-chicago keeps the IP and the content in the same place.
Do less work#
9. Use headless browsers only when needed#
Browsers are slow, expensive and expose more fingerprint surface. Before reaching for one, open your browser's network tab: many pages load their data from a JSON endpoint you can call directly. When JavaScript rendering is unavoidable, block images, fonts and media, and give each browser context its own sticky session. The Puppeteer and Playwright guide has the code.
10. Cache and deduplicate#
The cheapest request is the one you never send. Store what you fetch, skip URLs you already have, and use conditional requests (If-None-Match, If-Modified-Since) so unchanged pages come back as a small 304. Sitemaps and change feeds tell you what is new without recrawling everything. Keep the raw responses as well as the parsed data, so a parser bug means re-parsing, not re-crawling.
11. Treat CAPTCHAs as a signal, not an obstacle#
A CAPTCHA means the site has decided you look automated. Getting past it does not change that decision; it only hides the cause. Lower the rate, review headers and sessions, switch to a higher-trust pool if the IP is the problem, and pause the domain if challenges keep coming. Fewer challenges is the goal, not faster answers to them.
Measure everything#
12. Monitor the success rate per domain#
Track status codes, response sizes and latency per domain, and alert when the success rate drops. Check content, not just status codes: a 200 that contains a challenge page or an empty product grid is still a block. Common signals and how to respond:
| Signal | Usual meaning | Response |
|---|---|---|
429 or a Retry-After header | Rate limit reached | Slow down and honor the header |
403 on pages that used to work | IP reputation or fingerprint flagged | Review headers; move the domain to residential |
200 with a tiny body or challenge markup | Soft block | Count it as a failure and reduce the rate |
| Redirects to login or consent pages | Session or cookie problem | Use sticky sessions and keep cookies |
| Rising latency and timeouts | Target under load, or throttling you | Lower concurrency |
Gateway 502 | An upstream peer dropped | Retry; it is not a block |
Log the session ID or exit IP with every failure, so you can tell a bad IP from a bad request pattern. Watching these numbers turns blocking from a surprise into a metric you manage, and a falling success rate usually points to the technique above that needs attention before a domain stops working entirely. For architectures and pool choices by target type, see web scraping with proxies.
Key takeaways#
- Stay within robots.txt, terms of service and the law; polite scrapers trigger fewer defenses.
- Budget requests and concurrency per domain, and back off with jitter on 429 and 503.
- Rotate IPs from the right pool, and keep IP, cookies, headers and geography consistent within each flow.
- Prefer JSON endpoints and caches over headless browsers and repeated fetches.
- Treat CAPTCHAs and soft blocks as signals to slow down, and track the success rate per domain.
