You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Python网页爬虫脚本以提升爬取速度?

爬虫脚本优化请求:提升单日期赛事数据爬取速度

我写了一个Python脚本用于爬取赛事网站数据:输入赛事日期(格式为YYYY/MM/DD)和当日赛事数量后,脚本会爬取指定表格并将每场赛事数据保存为CSV文件。因为目标表格固定为12行,所以我在循环里硬编码了行数。

脚本可以正常运行,但爬取单日期数据需要20-30秒。我原本想用concurrent来提速,但没看到明显效果,希望能得到优化建议提升速度,非常感谢!

附原脚本代码:

import concurrent.futures
import pandas as pd
from requests_html import HTMLSession
import requests_cache

session = HTMLSession()
requests_cache.install_cache(expire_after=3600)

game_date = input("Please input the date of the game that you want to scrap (in YYYY/MM/DD): ")
game_no = int(input("Please input how many game on that date: "))

def split_list(big_list, chunk_size):
    return [big_list[i:i + chunk_size] for i in range(0, len(big_list), chunk_size)]

def get_game_result(game):
    print(f"Processing game {game}")
    url = f"https://example.com{game_date}&{game}"  <<<< example link
    response = session.get(url)
    response.html.render(sleep=5, keep_page=True, scrolldown=1)
    row_body = response.html.xpath(f"/html/body/div[1]/div[3]/div[2]/div[2]/div[2]/div[5]/table/tbody/tr[1]")

    final_list = []
    for i in range(2, 14):
        for item in row_body:
            item_table = item.text.split("\n")
            final_list.append(item_table)
            row_body = response.html.xpath(f"/html/body/div[1]/div[3]/div[2]/div[2]/div[2]/div[5]/table/tbody/tr[{i}]")
        i += 1

    df = pd.DataFrame(final_list)
    df.to_csv(f"game_{game}.csv", index=False, header=False)
    print(f"Finished processing game {game}")

with concurrent.futures.ThreadPoolExecutor(max_workers=10) as executor:
    results = [executor.submit(get_game_result, game) for game in range(1, game_no + 1)]
    for f in concurrent.futures.as_completed(results):
        f.result()

针对性优化建议(解决速度瓶颈)

  • 砍掉冗余的等待时间
    原脚本里render(sleep=5)是最大的耗时点,先确认网站是否真的需要等5秒才能加载表格。可以先把sleep降到1-2秒,甚至如果表格是服务器静态渲染的(查看网页源码能看到表格内容),直接去掉render()改用普通requests——requests_html的无头浏览器渲染本身就很重,能不用就不用。如果必须用无头浏览器,换成wait=2代替固定休眠,它会等待元素加载完成再继续,能节省无效等待时间。

  • 优化XPATH查询逻辑
    原代码循环里重复查询整个XPATH,逻辑还存在冗余:先取tr[1],再循环取tr[i],还重复添加数据。直接一次性获取所有12行tr元素,避免多次DOM查询:

    # 替换原循环逻辑
    rows = response.html.xpath("/html/body/div[1]/div[3]/div[2]/div[2]/div[2]/div[5]/table/tbody/tr[position()>=1 and position()<=12]")
    final_list = [row.text.split("\n") for row in rows]
    
  • 调整并发参数
    原代码max_workers=10可能太高,触发网站反爬限流反而变慢。可以降到5-6,同时检查requests_cache是否生效——在session.get()后打印response.from_cache,如果是True说明命中缓存,重复请求会快很多。

  • 替换更轻量的解析工具
    如果表格是静态渲染的,直接用requests+BeautifulSoup代替requests_html,速度会大幅提升,因为不需要启动无头浏览器。示例片段:

    import requests
    from bs4 import BeautifulSoup
    
    response = requests.get(url)
    soup = BeautifulSoup(response.text, 'html.parser')
    table_tbody = soup.select_one("div:nth-of-type(5) table tbody")
    rows = table_tbody.find_all('tr')[:12]
    final_list = [row.get_text(strip=False).split("\n") for row in rows]
    
  • 清理冗余代码
    原代码里的split_list函数定义了但没用到,循环里的i += 1也是多余的(range(2,14)已经自动递增),直接删掉即可。


内容的提问来源于stack exchange,提问作者Jacky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 07:35:16