Which proxy type should you use for scraping?
| Target site | Recommended proxy | Session mode |
|---|---|---|
| Public data, no bot protection | Datacenter | Fastest and cheapest per GB |
| E-commerce, travel, search engines | Rotating residential | New IP per request |
| Multi-step flows (search, then paginate) | Rotating residential | Sticky, 5–30 minutes |
| Logged-in sessions | ISP | Same IP for the whole job |
| Mobile-only or very strict sites | Mobile | Rotating or sticky |
How to avoid blocks when scraping
- Rotate IPs. With a residential pool, every request can exit from a different IP.
- Keep sessions consistent. If a site sets cookies on one page and checks them on the next, send both requests through the same sticky session.
- Limit concurrency per site. Hundreds of parallel requests to one domain look like an attack, whatever the IP.
- Send realistic headers. A current browser
User-Agent,Accept-Languageand normal header order. - Retry with backoff. Treat
403,429and timeouts as signals to slow down and switch IP. - Target the right country. Prices and content often change by region, so pick IPs in the market you are measuring.
Python example: rotating proxies with retries
import random
import time
import requests
PROXY = "http://USERNAME:PASSWORD@HOST:PORT" # values from your dashboard
proxies = {"http": PROXY, "https": PROXY}
headers = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/128.0 Safari/537.36",
"Accept-Language": "en-US,en;q=0.9",
}
def fetch(url: str, attempts: int = 4) -> str:
for attempt in range(attempts):
try:
response = requests.get(url, proxies=proxies, headers=headers, timeout=30)
if response.status_code == 200:
return response.text
except requests.RequestException:
pass
# exponential backoff with jitter; the next request gets a new IP
time.sleep(2**attempt + random.random())
raise RuntimeError(f"Failed after {attempts} attempts: {url}")
print(fetch("https://api.ipify.org"))Scraping responsibly
- Collect only public data you have a lawful basis to use.
- Respect
robots.txtand rate limits, and identify your scraper where appropriate. - Do not use proxies to log in to accounts you do not own or to bypass paywalls.
FuseProxy's acceptable use policy lists what is not allowed on the network.