如何为True People Search的Selenium爬虫集成代理并优化稳定性?
解决方案:Selenium集成代理+抗检测优化(针对True People Search)
1. 代理集成与随机选择
- 将400个代理整理为统一格式(带认证的用
http://username:password@ip:port,无认证用http://ip:port),建议存放在proxies.txt文件中(每行一个代理),避免硬编码。 - 用
random.choice()实现每次请求前随机抽取代理:
import random def load_proxies(proxy_file_path): with open(proxy_file_path, 'r') as f: return [line.strip() for line in f if line.strip()] # 加载代理池 proxies_pool = load_proxies('proxies.txt') # 随机获取一个代理 def get_random_proxy(): return random.choice(proxies_pool)
2. Selenium代理配置(Chrome为例)
通过ChromeOptions注入代理参数,同时配置基础反检测项:
from selenium import webdriver from selenium.webdriver.chrome.options import Options def init_chrome_with_proxy(proxy): chrome_options = Options() # 无头模式(可选,降低资源占用) chrome_options.add_argument('--headless=new') # 禁用图片加载,提升爬取速度 chrome_options.add_argument('--blink-settings=imagesEnabled=false') # 禁用GPU渲染 chrome_options.add_argument('--disable-gpu') # 模拟真实浏览器UA chrome_options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36') # 注入代理 chrome_options.add_argument(f'--proxy-server={proxy}') return webdriver.Chrome(options=chrome_options)
3. IP封禁与速率限制处理
通过异常捕获+指数退避重试机制,应对封禁和请求限制:
import time from selenium.common.exceptions import TimeoutException, WebDriverException MAX_RETRY_TIMES = 3 INITIAL_RETRY_DELAY = 2 def crawl_target_url(url): retry_count = 0 while retry_count < MAX_RETRY_TIMES: proxy = get_random_proxy() driver = None try: driver = init_chrome_with_proxy(proxy) driver.set_page_load_timeout(15) # 设置页面加载超时 driver.get(url) # 检测是否被封禁(根据True People Search的封禁页面特征调整) if "Access Denied" in driver.page_source or "captcha" in driver.page_source.lower(): raise WebDriverException("IP blocked or captcha triggered") # 这里写入你的数据爬取、存储逻辑 # ... driver.quit() return # 爬取成功,退出循环 except (TimeoutException, WebDriverException) as e: print(f"Proxy {proxy} failed: {str(e)}. Retry {retry_count+1}/{MAX_RETRY_TIMES}") retry_count += 1 time.sleep(INITIAL_RETRY_DELAY * (2 ** retry_count)) # 指数退避延迟 if driver: driver.quit() print(f"Failed to crawl {url} after {MAX_RETRY_TIMES} retries")
4. 高效可靠爬取的优化建议
- 代理池健康检查:定期过滤无效代理,避免浪费请求时间:
def validate_proxy(proxy): try: driver = init_chrome_with_proxy(proxy) driver.get("https://httpbin.org/ip") time.sleep(1) # 验证代理IP是否生效 if proxy.split('@')[-1].split(':')[0] in driver.page_source: driver.quit() return True driver.quit() return False except: return False # 过滤有效代理,更新代理池 proxies_pool = [p for p in proxies_pool if validate_proxy(p)]
- 模拟人类行为:在爬取请求之间添加1-5秒的随机延迟,避免固定频率触发检测
- 记录爬取状态:用CSV或本地文件记录已完成的爬取任务,避免重复请求
- 动态调整浏览器指纹:偶尔切换UA字符串、窗口大小,降低被识别为爬虫的概率
内容的提问来源于stack exchange,提问作者Maryam
相关产品推荐
相关产品推荐

