You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ThreadPoolExecutor结合Selenium遇重复Soup:为何数据重复数与线程数相关?

问题分析与解决思路

核心问题根源

你遇到的重复数据问题,本质是ChromeDriver实例不支持多线程并发访问:

  • 多个线程共用同一个driver对象时,线程间的driver.get(url)调用会互相覆盖——线程A刚发起页面请求,线程B就替换成了新的URL,最终所有线程拿到的都是最后完成加载的页面源码,导致SKU重复。
  • 另外,直接操作全局列表allItemsDetails虽在Python中append是原子操作,但多线程下操作全局变量并非最佳实践,存在潜在风险。

具体解决方案

方案1:每个线程创建独立ChromeDriver实例

适合URL数量不多的场景,每个线程拥有专属浏览器实例,彻底避免资源竞争:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup
import concurrent.futures

user_agent = "你的User-Agent字符串"

def getSoup(url):
    # 每个线程初始化独立的driver
    options = Options()
    options.add_argument("--headless")
    options.add_argument("--disable-extensions")
    options.add_argument("--disable-dev-shm-usage")
    options.add_argument("--no-sandbox")
    options.add_argument(f"user-agent={user_agent}")
    driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
    
    try:
        driver.get(url)
        return BeautifulSoup(driver.page_source, "lxml")
    finally:
        # 用完立即关闭,释放资源
        driver.quit()

def getItemDetails(itemUrl):
    try:
        soup = getSoup(itemUrl)
        sku = soup.find("span", {"itemprop": "sku"}).text
        print(sku)
        return {"sku": sku}
    except Exception as e:
        print(f"处理出错: {str(e)}")
        return None

if __name__ == "__main__":
    allItemUrls = ["你的URL1", "你的URL2", "你的URL3", "你的URL4"]
    allItemsDetails = []
    
    with concurrent.futures.ThreadPoolExecutor(max_workers=4) as executor:
        # 收集任务返回值,替代全局变量操作
        results = executor.map(getItemDetails, allItemUrls)
        allItemsDetails = [res for res in results if res is not None]
    
    print(allItemsDetails)

方案2:改用requests库(静态页面优先)

如果目标页面是静态HTML、无需JS渲染,用requests比Selenium更高效,且天然线程安全:

import requests
from bs4 import BeautifulSoup
import concurrent.futures

user_agent = "你的User-Agent字符串"
headers = {"User-Agent": user_agent}

def getSoup(url):
    response = requests.get(url, headers=headers)
    response.raise_for_status()  # 抛出HTTP请求错误
    return BeautifulSoup(response.text, "lxml")

def getItemDetails(itemUrl):
    try:
        soup = getSoup(itemUrl)
        sku = soup.find("span", {"itemprop": "sku"}).text
        print(sku)
        return {"sku": sku}
    except Exception as e:
        print(f"处理出错: {str(e)}")
        return None

if __name__ == "__main__":
    allItemUrls = ["你的URL1", "你的URL2", "你的URL3", "你的URL4"]
    allItemsDetails = []
    
    with concurrent.futures.ThreadPoolExecutor(max_workers=4) as executor:
        results = executor.map(getItemDetails, allItemUrls)
        allItemsDetails = [res for res in results if res is not None]
    
    print(allItemsDetails)

方案3:用多进程替代多线程(大量URL场景)

ChromeDriver是进程级资源,多进程可天然隔离每个driver实例,适合URL数量较多的情况:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup
import concurrent.futures

user_agent = "你的User-Agent字符串"

def getSoup(url):
    options = Options()
    options.add_argument("--headless")
    options.add_argument("--disable-extensions")
    options.add_argument("--disable-dev-shm-usage")
    options.add_argument("--no-sandbox")
    options.add_argument(f"user-agent={user_agent}")
    driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options)
    
    try:
        driver.get(url)
        return BeautifulSoup(driver.page_source, "lxml")
    finally:
        driver.quit()

def getItemDetails(itemUrl):
    try:
        soup = getSoup(itemUrl)
        sku = soup.find("span", {"itemprop": "sku"}).text
        print(sku)
        return {"sku": sku}
    except Exception as e:
        print(f"处理出错: {str(e)}")
        return None

if __name__ == "__main__":
    allItemUrls = ["你的URL1", "你的URL2", "你的URL3", "你的URL4"]
    allItemsDetails = []
    
    # 替换为ProcessPoolExecutor
    with concurrent.futures.ProcessPoolExecutor(max_workers=4) as executor:
        results = executor.map(getItemDetails, allItemUrls)
        allItemsDetails = [res for res in results if res is not None]
    
    print(allItemsDetails)

内容的提问来源于stack exchange,提问作者bltz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 09:01:03