You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为True People Search的Selenium爬虫集成代理并优化稳定性?

解决方案:Selenium集成代理+抗检测优化(针对True People Search)

1. 代理集成与随机选择

  • 将400个代理整理为统一格式(带认证的用http://username:password@ip:port,无认证用http://ip:port),建议存放在proxies.txt文件中(每行一个代理),避免硬编码。
  • 用random.choice()实现每次请求前随机抽取代理:
import random

def load_proxies(proxy_file_path):
    with open(proxy_file_path, 'r') as f:
        return [line.strip() for line in f if line.strip()]

# 加载代理池
proxies_pool = load_proxies('proxies.txt')

# 随机获取一个代理
def get_random_proxy():
    return random.choice(proxies_pool)

2. Selenium代理配置(Chrome为例)

通过ChromeOptions注入代理参数,同时配置基础反检测项:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

def init_chrome_with_proxy(proxy):
    chrome_options = Options()
    # 无头模式(可选,降低资源占用)
    chrome_options.add_argument('--headless=new')
    # 禁用图片加载,提升爬取速度
    chrome_options.add_argument('--blink-settings=imagesEnabled=false')
    # 禁用GPU渲染
    chrome_options.add_argument('--disable-gpu')
    # 模拟真实浏览器UA
    chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36')
    # 注入代理
    chrome_options.add_argument(f'--proxy-server={proxy}')
    
    return webdriver.Chrome(options=chrome_options)

3. IP封禁与速率限制处理

通过异常捕获+指数退避重试机制,应对封禁和请求限制:

import time
from selenium.common.exceptions import TimeoutException, WebDriverException

MAX_RETRY_TIMES = 3
INITIAL_RETRY_DELAY = 2

def crawl_target_url(url):
    retry_count = 0
    while retry_count < MAX_RETRY_TIMES:
        proxy = get_random_proxy()
        driver = None
        try:
            driver = init_chrome_with_proxy(proxy)
            driver.set_page_load_timeout(15)  # 设置页面加载超时
            driver.get(url)
            
            # 检测是否被封禁(根据True People Search的封禁页面特征调整)
            if "Access Denied" in driver.page_source or "captcha" in driver.page_source.lower():
                raise WebDriverException("IP blocked or captcha triggered")
            
            # 这里写入你的数据爬取、存储逻辑
            # ...
            
            driver.quit()
            return  # 爬取成功,退出循环
        except (TimeoutException, WebDriverException) as e:
            print(f"Proxy {proxy} failed: {str(e)}. Retry {retry_count+1}/{MAX_RETRY_TIMES}")
            retry_count += 1
            time.sleep(INITIAL_RETRY_DELAY * (2 ** retry_count))  # 指数退避延迟
            if driver:
                driver.quit()
    print(f"Failed to crawl {url} after {MAX_RETRY_TIMES} retries")

4. 高效可靠爬取的优化建议

  • 代理池健康检查:定期过滤无效代理,避免浪费请求时间:
def validate_proxy(proxy):
    try:
        driver = init_chrome_with_proxy(proxy)
        driver.get("https://httpbin.org/ip")
        time.sleep(1)
        # 验证代理IP是否生效
        if proxy.split('@')[-1].split(':')[0] in driver.page_source:
            driver.quit()
            return True
        driver.quit()
        return False
    except:
        return False

# 过滤有效代理,更新代理池
proxies_pool = [p for p in proxies_pool if validate_proxy(p)]
  • 模拟人类行为:在爬取请求之间添加1-5秒的随机延迟,避免固定频率触发检测
  • 记录爬取状态:用CSV或本地文件记录已完成的爬取任务,避免重复请求
  • 动态调整浏览器指纹:偶尔切换UA字符串、窗口大小,降低被识别为爬虫的概率

内容的提问来源于stack exchange,提问作者Maryam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 06:13:20