You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为网页抓取实现多线程以加速Python程序运行?

多线程优化网页抓取程序方案

先修正原函数的错误

你原代码里的data = json.loads(data)存在未定义变量的问题,直接运行会报错。假设JSON数据是从页面的application/json类型script标签提取,修正后的函数如下:

import requests
from bs4 import BeautifulSoup
import json
import pandas as pd
from concurrent.futures import ThreadPoolExecutor

def webscrape(url):
    # 加超时防止请求卡住
    response = requests.get(url, timeout=10)
    # 可选:加请求头模拟浏览器,避免反爬
    # headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}
    # response = requests.get(url, headers=headers, timeout=10)
    soup = BeautifulSoup(response.content, "html.parser")
    # 根据实际页面结构调整JSON数据的提取方式
    json_script = soup.find("script", type="application/json")
    if not json_script:
        return pd.DataFrame()
    data = json.loads(json_script.string)
    return pd.DataFrame(data['dataProvider'])

多线程实现代码

用concurrent.futures.ThreadPoolExecutor实现并行抓取,比串行效率高很多:

# 替换成你的实际URL列表
urls = ["https://example.com/page1", "https://example.com/page2", "https://example.com/page3"]

# 创建线程池,max_workers根据需求调整,建议5-10
with ThreadPoolExecutor(max_workers=8) as executor:
    # 批量提交任务,按URL顺序返回结果
    results = list(executor.map(webscrape, urls))

# 过滤空DataFrame后合并
final_df = pd.concat([df for df in results if not df.empty], axis=1)

关键注意事项

  • 线程数设置:max_workers不要设得过大,否则容易触发目标网站的反爬机制,导致IP被限制。
  • 异常处理:可以在webscrape函数里加try-except块,捕获请求或解析错误,避免单个任务失败影响整个线程池。
  • 反爬应对:添加合理的请求头、适当的延迟(如果需要),避免给目标服务器造成过大压力。

内容的提问来源于stack exchange,提问作者Jacques

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 02:45:07