You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

优化基于concurrent.futures的Python网页爬虫脚本以缩短执行时间

爬虫脚本性能优化方案

针对你的爬虫脚本耗时395秒的问题,以下是具体的优化方向和实现建议,涵盖并发模型、HTTP请求、数据处理等多个维度:

一、并发模型替换:从ProcessPoolExecutor改为ThreadPoolExecutor(或异步库)

你的爬虫属于IO密集型任务(大部分时间在等待HTTP响应),ProcessPoolExecutor的进程创建、上下文切换和IPC开销远高于ThreadPoolExecutor,是当前性能瓶颈的核心原因之一。

具体调整:

  1. 替换ProcessPoolExecutor为ThreadPoolExecutor,并适当调高max_workers(建议10-15,需根据网站反爬策略调整,避免触发429):
    from concurrent.futures import ThreadPoolExecutor  # 替换导入
    # ...
    with ThreadPoolExecutor(max_workers=12) as executor:
        # 后续逻辑保持一致(或优化任务提交方式)
    
  2. 进阶优化:改用异步HTTP库aiohttp配合asyncio,异步非阻塞模型能同时处理更多请求,性能比线程池提升更显著。示例框架:
    import aiohttp
    import asyncio
    
    async def get_page_content_async(session, url):
        headers = {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.51 Safari/537.36',
            'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
            'Accept-Language': 'en-US,en;q=0.5',
            'Accept-Encoding': 'gzip, deflate, br',
            'Connection': 'keep-alive',
            'Upgrade-Insecure-Requests': '1'
        }
        async with session.get(url, headers=headers) as response:
            return await response.text()  # aiohttp自动处理gzip/br解码
    

二、HTTP请求层优化

1. 改用requests库替代urllib,并复用会话

requests内置了自动解码gzip/brotli的功能,无需手动处理解压逻辑,同时Session对象可以复用TCP连接,大幅减少TCP三次握手的开销:

import requests

# 全局创建Session,复用连接
session = requests.Session()
session.headers.update({
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.51 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Accept-Encoding': 'gzip, deflate, br',
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
})

def get_page_content(url):
    response = session.get(url)
    response.raise_for_status()  # 自动抛出HTTP错误
    return response.text  # 自动解码

2. 优化重试机制

  • 优先读取响应头的Retry-After字段(429错误时返回),代替固定延迟:
    except requests.exceptions.HTTPError as e:
        if e.response.status_code == 429:
            retry_after = int(e.response.headers.get('Retry-After', retry_delay * retries))
            time.sleep(retry_after)
    
  • 用tenacity库简化重试逻辑,代码更简洁且可配置性更强:
    from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
    
    @retry(stop=stop_after_attempt(max_retries),
           wait=wait_exponential(multiplier=1, min=2, max=10),
           retry=retry_if_exception_type((requests.exceptions.HTTPError, requests.exceptions.ConnectionError)))
    def extract_table_data(page_url, page_number):
        # 原有逻辑
    

三、数据处理优化

1. 批量拼接DataFrame,避免多次pd.concat

当前每次获取到df就执行pd.concat,会频繁创建新的DataFrame,内存开销大且效率低。改为先收集所有df到列表,最后一次性拼接:

all_dfs = []  # 替换原来的all_data = pd.DataFrame()
# ...
# 在处理结果时:
if df is not None:
    all_dfs.append(df)
    consecutive_errors = 0
# ...
# 最后统一拼接
all_data = pd.concat(all_dfs, ignore_index=True)

2. 提升HTML解析速度

将BeautifulSoup的解析器从html.parser改为lxml(需先安装pip install lxml),解析速度提升数倍:

soup = BeautifulSoup(webpage, 'lxml')

3. 简化表格提取逻辑

pd.read_html可以直接接受HTML文本,无需转成StringIO,减少中间步骤:

df = pd.read_html(webpage)[0]

四、其他细节优化

1. 提前获取总页数,批量提交任务

当前动态添加任务的循环逻辑存在冗余,可先爬取第一页解析出总页数,然后一次性提交所有页面的任务,减少循环判断开销:

# 先获取总页数
def get_total_pages():
    url = base_url + '1'
    webpage = get_page_content(url)
    soup = BeautifulSoup(webpage, 'lxml')
    # 解析分页栏的总页数,示例(需根据实际页面结构调整)
    total_page_elem = soup.find('li', class_='page-item last')
    return int(total_page_elem.text.strip())

total_pages = get_total_pages()
# 批量提交任务
with ThreadPoolExecutor(max_workers=12) as executor:
    futures = {executor.submit(process_page, page): page for page in range(1, total_pages+1)}
    # 处理结果...

2. 优化日志与终止逻辑

  • 移除log_stream的依赖,改用变量跟踪404状态,避免频繁读写StringIO:
    encountered_404 = False
    # ...
    except urllib.error.HTTPError as e:
        if e.code == 404:
            logger.info("Reached a page that does not exist. Stopping.")
            encountered_404 = True
            break
    # ...
    if consecutive_errors >= max_consecutive_errors or encountered_404:
        break
    
  • 减少不必要的traceback.print_exc()调用,仅在调试阶段使用,生产环境记录关键错误信息即可。

3. 调整并发数与延迟

根据网站的反爬策略,逐步调高max_workers(从10开始测试),同时在请求之间添加随机小延迟(0.5-2秒),避免被识别为爬虫:

import random

def process_page(page):
    # ...
    time.sleep(random.uniform(0.5, 2))  # 在请求前/后添加随机延迟

效果预期

通过以上优化,尤其是并发模型替换、会话复用和批量数据处理,脚本执行时间可缩短至原来的1/3甚至更低(具体取决于页面数量和网站响应速度)。

内容的提问来源于stack exchange,提问作者HamidBee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 14:50:06