You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用ThreadPoolExecutor结合Excel实现网站邮编数据爬取与写入

优化方案:用ThreadPoolExecutor提升邮编数据爬取效率

你的单线程脚本可以正常运行,但多线程实现需要调整——openpyxl的工作表和单元格对象并非线程安全,直接在子线程中修改单元格值可能引发未知问题。以下是修正后的多线程实现:

修正后的代码

import concurrent.futures as futures
from openpyxl import load_workbook
from bs4 import BeautifulSoup
import requests

headers = {
    'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/109.0.0.0 Safari/537.36',
    'accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9',
    'accept-encoding': 'gzip, deflate, br',
    'accept-language': 'en-US,en;q=0.9,bn;q=0.8',
    'referer': 'https://www.unitedstateszipcodes.org/60634/'
}

link = 'https://www.unitedstateszipcodes.org/{}/'

# 仅负责爬取数据,不直接操作Excel
def get_content(zip_code, url):
    try:
        res = requests.get(url, headers=headers)
        res.raise_for_status()  # 捕获HTTP请求错误
        soup = BeautifulSoup(res.text, "lxml")
        household_text = soup.select_one("th:-soup-contains('Occupied Housing Units') + td").get_text(strip=True)
        household = int(household_text.replace(",", ""))
        return (zip_code, household)
    except Exception as e:
        print(f"处理邮编 {zip_code} 时出错: {str(e)}")
        return (zip_code, None)

if __name__ == '__main__':
    wb = load_workbook('input.xlsx')
    ws = wb['result']
    
    # 收集需要处理的邮编及对应行号
    zip_data = []
    for row_idx, cell in enumerate(ws['B'], start=1):
        if cell.value == "Zip Code":
            continue
        if not cell.value:
            break
        zip_code = cell.value
        zip_data.append((zip_code, row_idx))
    
    # 多线程批量爬取
    results = []
    with futures.ThreadPoolExecutor(max_workers=6) as executor:
        future_to_zip = {
            executor.submit(get_content, zip_code, link.format(zip_code)): (zip_code, row_idx) 
            for zip_code, row_idx in zip_data
        }
        # 处理完成的任务
        for future in futures.as_completed(future_to_zip):
            zip_code, row_idx = future_to_zip[future]
            try:
                _, household = future.result()
                results.append((row_idx, household))
            except Exception as e:
                print(f"获取邮编 {zip_code} 结果时出错: {str(e)}")
                results.append((row_idx, None))
    
    # 主线程统一写入Excel
    for row_idx, household in results:
        if household is not None:
            ws.cell(row=row_idx, column=4).value = household  # B列偏移3列对应D列
    
    wb.save('input.xlsx')
    print("数据处理完成,已保存到文件")

关键修改点

  • 线程安全优化:子线程仅返回爬取结果,由主线程统一写入Excel,避免多线程操作openpyxl对象的风险
  • 错误容错:添加异常捕获,单个邮编处理失败不会中断整个程序
  • 精准映射:提前记录每个邮编对应的行号,确保结果写入正确单元格
  • 资源自动管理:使用线程池上下文管理器,自动回收线程资源

内容的提问来源于stack exchange,提问作者robots.txt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 19:41:05