You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python批量发送GET请求爬取房产网站数据遇代理及反爬问题求助

Fixing Proxy Connection Errors & Anti-Bot Blocks for Meilleurs-Agents Scraper

Got it, let's work through your two main issues here: failing proxy connections, and getting flagged as a bot when you skip proxies. Here's how to resolve both:

1. Fixing the Proxy Connection Failure

The error getaddrinfo failed usually points to a problem with your proxy setup. Let's troubleshoot this first:

  • Add the proxy protocol prefix: Requests requires you to specify the protocol for the proxy (even if your target URL uses HTTPS). Most people forget this—update your proxy definition to include http:// before the IP address:

    PROXY = {'https': 'http://XX.XXX.X.XXX:XXXX'}  # Note the http:// prefix
    

    This tells requests how to properly communicate with the proxy server.

  • Test the proxy independently: Before integrating it into your scraper, verify the proxy is alive and works with HTTPS requests. Run this quick test script:

    import requests
    test_proxy = {'https': 'http://XX.XXX.X.XXX:XXXX'}
    try:
        # Use httpbin.org to check if the proxy routes your request correctly
        resp = requests.get('https://httpbin.org/ip', proxies=test_proxy, timeout=10)
        print("Proxy working! Your IP via proxy:", resp.json())
    except Exception as e:
        print(f"Proxy test failed: {str(e)}")
    

    If this fails, your proxy is either dead, requires authentication, or is blocked by your network.

  • Check for proxy authentication: If your proxy needs a username/password, format it like this:

    PROXY = {'https': 'http://username:password@XX.XXX.X.XXX:XXXX'}
    

    Free proxies often stop working quickly, so consider using a paid proxy service if you need reliable, long-term access.

2. Bypassing the Anti-Bot Detection

Even with a working proxy, you'll still get blocked if your requests don't mimic a real browser. Here's how to make your scraper look legitimate:

  • Update your request headers: Your current User-Agent is an old Firefox version—swap it for a modern one, and add other common headers that real browsers send:

    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept-Language': 'fr-FR,fr;q=0.9,en-US;q=0.8,en;q=0.7',  # Match the site's language
        'Accept-Encoding': 'gzip, deflate, br',
        'Referer': 'https://www.meilleursagents.com/',
        'DNT': '1'  # Do Not Track header, common in real browsers
    }
    
  • Use a session to maintain state: A requests.Session() preserves cookies and headers across requests, just like a real browser. This helps avoid being flagged as a bot:

    import requests
    import time
    import random
    
    session = requests.Session()
    session.headers.update(headers)
    session.proxies.update(PROXY)
    
    for city, postal_code in zip(cities, postal_codes):
        url = f'https://www.meilleursagents.com/prix-immobilier/{city}-{postal_code}/'
        try:
            response = session.get(url, timeout=10)
            # Check if you're still blocked
            if "you seems to be a bot" in response.text:
                print(f"Blocked on {city}! Adjust delays or headers.")
            else:
                # Process your target data here
                print(f"Successfully fetched {city} data")
            # Add random delay to mimic human browsing patterns
            time.sleep(random.uniform(2, 5))
        except Exception as e:
            print(f"Error fetching {city}: {str(e)}")
    
  • Add random delays: Spamming requests in a tight loop is a dead giveaway for bots. Adding random pauses between requests (2-5 seconds) makes your scraper behave more like a real user navigating the site.

  • Consider a headless browser (if needed): If the site uses JavaScript to load data or detect bots, requests might not be enough. Tools like Playwright or Selenium can simulate a full browser, including scrolling and clicking, which makes it much harder to detect your scraper.

Final Notes

Always make sure you're complying with the site's robots.txt and terms of service—scraping can violate some websites' policies. Start with slow request rates and test small batches before scaling up.

内容的提问来源于stack exchange,提问作者Nelly Barret

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:36:17