You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬虫竞态问题排查:重复URL与年份匹配错误解决

多线程爬虫年份与PDF链接不匹配的问题修复

问题根源

  • WebDriver线程不安全:你在scrape_data中创建了单个ChromeDriver实例,多个线程共享该实例。当线程A刚调用driver.get(2015年URL),线程B立刻执行driver.get(2022年URL),会导致线程A后续解析的页面是2022年的内容,但仍然用2015年的年份标记,最终出现年份和链接不匹配的错误。
  • 共享WebDriver引发的竞态条件,完全破坏了多线程的独立性,导致数据关联错误。

修复方案

  1. 每个线程独立创建WebDriver:在scrape_year函数内部初始化ChromeDriver,确保每个线程拥有自己的浏览器实例,避免资源竞争。
  2. 移除共享的data参数:让scrape_year独立返回当前年份的结果,再由主线程汇总,避免线程直接操作共享列表。
  3. 修复BeautifulSoup元素查找逻辑:原代码中字典重复的class键会被覆盖,改用列表形式传入目标class,确保所有需要的元素都能被匹配。

修改后的完整代码

import os
from apify_client import ApifyClient
import concurrent.futures
from selenium.common.exceptions import TimeoutException
import requests
import subprocess
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import boto3
from datetime import datetime
from selenium.webdriver.chrome.service import Service
import stat
import json
import ast
import threading
import time
from concurrent.futures import ThreadPoolExecutor


def main():
    s3 = boto3.client('s3')
    today_date = datetime.today().strftime('%Y-%m-%d')
    generate_xmlfiles(s3, today_date)
    print("Uploaded successfully")

def generate_xmlfiles(s3, today_date):
    s3.put_object(Bucket='data', Key='Bany/'+today_date+"/xmlfiles.txt", Body=str(scrape_data()))
    print("Exported files to s3")


def scrape_data():  
    years = [2022, 2015]
    data = []
    states = ["alnb"]

    with requests.Session() as session:
        for state in states:
            with ThreadPoolExecutor() as executor:
                futures = []
                for year in years:
                    url = f'https://www.govinfo.gov/app/collection/uscourts/bankruptcy/{state}/{year}/%7B%22pageSize%22%3A%22100%22%2C%22offset%22%3A%220%22%7D'
                    response = session.get(url)
                    if response.status_code == 200:               
                        # 每个线程独立初始化driver,不再共享资源
                        futures.append(executor.submit(scrape_year, url, state, year))  
                for future in futures:
                    data += future.result()
    
            print("Loaded " + state)
        
    return data
    
def scrape_year(url, state, year):
    print(f"scraping data for state {state.capitalize()} for {str(year)}")
    # 每个线程创建独立的WebDriver实例
    options = webdriver.ChromeOptions()
    options.add_argument("--no-sandbox")
    options.add_argument("--disable-dev-shm-usage")
    options.add_argument("headless")
    driver = webdriver.Chrome(options=options)
    
    try:
        driver.get(url)
        driver.implicitly_wait(10)
        try_count = 0
        while try_count < 3:
            try:
                WebDriverWait(driver, 120).until(EC.presence_of_all_elements_located((By.CLASS_NAME, "panel-body")))
                soup = BeautifulSoup(driver.page_source, 'html.parser')
                # 修复class参数:用列表替代重复key的字典
                bankruptcy_element = soup.findAll('div', class_=["panel-collapse collapse in", "panel-title", "panel-body", "panel panel-default", "btn-group-horizontal"])          
                return [{year: f"https://www.govinfo.gov/metadata/granule{xmlfile['href'].replace('.pdf','/mods.xml').replace('/pdf','').replace('/pkg/','/').replace('/content','')}"} 
                        for i in bankruptcy_element 
                        for xmlfile in i.findAll('a', href=True) 
                        if "pdf" in xmlfile['href']]
            except TimeoutException:
                print(f"TimeoutException encountered. Retrying {try_count + 1} of 3...")
                try_count += 1
    finally:
        # 确保每个线程的driver都被关闭,避免进程泄漏
        driver.quit()
     
main()

额外优化说明

  • 使用f-string替换字符串拼接,提升代码可读性。
  • 通过finally块强制关闭每个线程的WebDriver,避免浏览器进程残留。

内容的提问来源于stack exchange,提问作者Xi12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 04:27:16