You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于SeleniumBase多线程+不同代理的爬虫IP重复问题求助

问题分析

  1. 流程不符合需求:原代码在单个页面爬取完成后立即串行处理该页面的商品链接,没有先收集所有页面的商品链接再统一并行爬取,违反了任务描述的要求。
  2. 代理生效异常:虽然代码尝试为不同页面/商品Driver分配不同代理,但可能因undetected_chromedriver代理参数传递逻辑、线程上下文问题,导致代理未正确独立生效。

解决方案

1. 重构为两阶段执行流程

  • 阶段一:并行爬取页面,批量收集商品链接
    启动多线程,每个线程用独立代理爬取指定页面,收集链接后立即销毁Driver,最后汇总所有商品链接。
  • 阶段二:并行爬取商品详情,每个Driver绑定唯一代理
    汇总所有链接后,启动新线程池,每个线程创建带独立代理的Driver,并行处理商品详情。

2. 确保代理正确传递

确认undetected_chromedriver的代理参数为完整格式,例如http://用户名:密码@IP:端口或socks5://用户名:密码@IP:端口。

修改后的代码

import concurrent.futures
import sys
from selenium.webdriver import ActionChains
import time
from selenium.webdriver.common.keys import Keys
from undetected_chromedriver import Driver  # 确保导入正确的Driver类

# 为undetected_chromedriver添加必要启动参数
sys.argv.append("-n")

pages = 3
# 替换为实际可用的代理,需符合完整格式
pool_proxies_for_pages = ['proxy0', 'proxy1', 'proxy2']
pool_proxies_for_products = ['proxy5', 'proxy6', 'proxy7']

def create_undetected_webdriver(proxy):
    """创建带指定代理的undetected Chrome Driver"""
    driver = Driver(
        uc=True,
        proxy=proxy,
        user_agent='Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36'
    )
    return driver

def parse_page(page_idx, proxy):
    """爬取单个页面,返回该页面的有效商品链接"""
    driver = None
    try:
        driver = create_undetected_webdriver(proxy)
        # eBay页码从1开始,修正原代码的页码偏移问题
        url = f'https://www.ebay.com/e/_electronics/shop-all-ebay-refurbished-cell-phones?_pgn={page_idx + 1}'
        driver.get(url)
        
        # 滚动加载页面内容
        ActionChains(driver).send_keys(Keys.END).perform()
        time.sleep(7)
        
        # 提取有效商品链接,过滤空值
        elements = driver.find_elements("xpath", "//a[@tabindex='-1']")
        links = [elem.get_attribute("href") for elem in elements if elem.get_attribute("href")]
        return links
    finally:
        # 确保Driver被销毁,避免资源泄漏
        if driver:
            driver.quit()

def parse_product(product_link, proxy):
    """爬取单个商品详情"""
    driver = None
    try:
        driver = create_undetected_webdriver(proxy)
        driver.get(product_link)
        # 此处添加商品详情提取逻辑,例如获取标题、价格等
        # 示例:title = driver.find_element("xpath", "//h1[@class='x-item-title__mainTitle']").text
        print(f"完成爬取: {product_link}")
    finally:
        if driver:
            driver.quit()

if __name__ == "__main__":
    # 阶段一:并行爬取所有页面,收集商品链接
    all_product_links = []
    with concurrent.futures.ThreadPoolExecutor(max_workers=pages) as page_executor:
        # 提交所有页面爬取任务
        futures = [page_executor.submit(parse_page, idx, proxy) for idx, proxy in enumerate(pool_proxies_for_pages)]
        # 等待任务完成,汇总所有链接
        for future in concurrent.futures.as_completed(futures):
            links = future.result()
            if links:
                all_product_links.extend(links)
    
    print(f"共收集到 {len(all_product_links)} 个商品链接")
    
    # 阶段二:并行爬取商品详情,循环分配代理池中的代理
    with concurrent.futures.ThreadPoolExecutor(max_workers=len(pool_proxies_for_products)) as product_executor:
        for link_idx, link in enumerate(all_product_links):
            # 循环使用代理池,确保每个Driver绑定唯一代理
            proxy = pool_proxies_for_products[link_idx % len(pool_proxies_for_products)]
            product_executor.submit(parse_product, link, proxy)

关键改进点

  • 分阶段执行:严格遵循任务要求,先完成所有页面的链接收集,再统一并行处理商品详情。
  • 代理循环分配:当商品数量多于代理池数量时,循环复用代理,确保每个Driver使用独立代理。
  • 资源安全回收:通过try...finally确保Driver无论是否异常都会被销毁,避免内存泄漏。
  • 页码修正:修复原代码中eBay页码从0开始的错误,保证爬取正确页面。

内容的提问来源于stack exchange,提问作者Nico

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 11:03:12