Python批量发送GET请求爬取房产网站数据遇代理及反爬问题求助
Got it, let's work through your two main issues here: failing proxy connections, and getting flagged as a bot when you skip proxies. Here's how to resolve both:
1. Fixing the Proxy Connection Failure
The error getaddrinfo failed usually points to a problem with your proxy setup. Let's troubleshoot this first:
Add the proxy protocol prefix: Requests requires you to specify the protocol for the proxy (even if your target URL uses HTTPS). Most people forget this—update your proxy definition to include
http://before the IP address:PROXY = {'https': 'http://XX.XXX.X.XXX:XXXX'} # Note the http:// prefixThis tells requests how to properly communicate with the proxy server.
Test the proxy independently: Before integrating it into your scraper, verify the proxy is alive and works with HTTPS requests. Run this quick test script:
import requests test_proxy = {'https': 'http://XX.XXX.X.XXX:XXXX'} try: # Use httpbin.org to check if the proxy routes your request correctly resp = requests.get('https://httpbin.org/ip', proxies=test_proxy, timeout=10) print("Proxy working! Your IP via proxy:", resp.json()) except Exception as e: print(f"Proxy test failed: {str(e)}")If this fails, your proxy is either dead, requires authentication, or is blocked by your network.
Check for proxy authentication: If your proxy needs a username/password, format it like this:
PROXY = {'https': 'http://username:password@XX.XXX.X.XXX:XXXX'}Free proxies often stop working quickly, so consider using a paid proxy service if you need reliable, long-term access.
2. Bypassing the Anti-Bot Detection
Even with a working proxy, you'll still get blocked if your requests don't mimic a real browser. Here's how to make your scraper look legitimate:
Update your request headers: Your current User-Agent is an old Firefox version—swap it for a modern one, and add other common headers that real browsers send:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'fr-FR,fr;q=0.9,en-US;q=0.8,en;q=0.7', # Match the site's language 'Accept-Encoding': 'gzip, deflate, br', 'Referer': 'https://www.meilleursagents.com/', 'DNT': '1' # Do Not Track header, common in real browsers }Use a session to maintain state: A
requests.Session()preserves cookies and headers across requests, just like a real browser. This helps avoid being flagged as a bot:import requests import time import random session = requests.Session() session.headers.update(headers) session.proxies.update(PROXY) for city, postal_code in zip(cities, postal_codes): url = f'https://www.meilleursagents.com/prix-immobilier/{city}-{postal_code}/' try: response = session.get(url, timeout=10) # Check if you're still blocked if "you seems to be a bot" in response.text: print(f"Blocked on {city}! Adjust delays or headers.") else: # Process your target data here print(f"Successfully fetched {city} data") # Add random delay to mimic human browsing patterns time.sleep(random.uniform(2, 5)) except Exception as e: print(f"Error fetching {city}: {str(e)}")Add random delays: Spamming requests in a tight loop is a dead giveaway for bots. Adding random pauses between requests (2-5 seconds) makes your scraper behave more like a real user navigating the site.
Consider a headless browser (if needed): If the site uses JavaScript to load data or detect bots,
requestsmight not be enough. Tools like Playwright or Selenium can simulate a full browser, including scrolling and clicking, which makes it much harder to detect your scraper.
Final Notes
Always make sure you're complying with the site's robots.txt and terms of service—scraping can violate some websites' policies. Start with slow request rates and test small batches before scaling up.
内容的提问来源于stack exchange,提问作者Nelly Barret

