如何优化Python网页爬虫脚本以提升爬取速度?
我写了一个Python脚本用于爬取赛事网站数据:输入赛事日期(格式为YYYY/MM/DD)和当日赛事数量后,脚本会爬取指定表格并将每场赛事数据保存为CSV文件。因为目标表格固定为12行,所以我在循环里硬编码了行数。
脚本可以正常运行,但爬取单日期数据需要20-30秒。我原本想用concurrent来提速,但没看到明显效果,希望能得到优化建议提升速度,非常感谢!
附原脚本代码:
import concurrent.futures import pandas as pd from requests_html import HTMLSession import requests_cache session = HTMLSession() requests_cache.install_cache(expire_after=3600) game_date = input("Please input the date of the game that you want to scrap (in YYYY/MM/DD): ") game_no = int(input("Please input how many game on that date: ")) def split_list(big_list, chunk_size): return [big_list[i:i + chunk_size] for i in range(0, len(big_list), chunk_size)] def get_game_result(game): print(f"Processing game {game}") url = f"https://example.com{game_date}&{game}" <<<< example link response = session.get(url) response.html.render(sleep=5, keep_page=True, scrolldown=1) row_body = response.html.xpath(f"/html/body/div[1]/div[3]/div[2]/div[2]/div[2]/div[5]/table/tbody/tr[1]") final_list = [] for i in range(2, 14): for item in row_body: item_table = item.text.split("\n") final_list.append(item_table) row_body = response.html.xpath(f"/html/body/div[1]/div[3]/div[2]/div[2]/div[2]/div[5]/table/tbody/tr[{i}]") i += 1 df = pd.DataFrame(final_list) df.to_csv(f"game_{game}.csv", index=False, header=False) print(f"Finished processing game {game}") with concurrent.futures.ThreadPoolExecutor(max_workers=10) as executor: results = [executor.submit(get_game_result, game) for game in range(1, game_no + 1)] for f in concurrent.futures.as_completed(results): f.result()
针对性优化建议(解决速度瓶颈)
砍掉冗余的等待时间
原脚本里render(sleep=5)是最大的耗时点,先确认网站是否真的需要等5秒才能加载表格。可以先把sleep降到1-2秒,甚至如果表格是服务器静态渲染的(查看网页源码能看到表格内容),直接去掉render()改用普通requests——requests_html的无头浏览器渲染本身就很重,能不用就不用。如果必须用无头浏览器,换成wait=2代替固定休眠,它会等待元素加载完成再继续,能节省无效等待时间。优化XPATH查询逻辑
原代码循环里重复查询整个XPATH,逻辑还存在冗余:先取tr[1],再循环取tr[i],还重复添加数据。直接一次性获取所有12行tr元素,避免多次DOM查询:# 替换原循环逻辑 rows = response.html.xpath("/html/body/div[1]/div[3]/div[2]/div[2]/div[2]/div[5]/table/tbody/tr[position()>=1 and position()<=12]") final_list = [row.text.split("\n") for row in rows]调整并发参数
原代码max_workers=10可能太高,触发网站反爬限流反而变慢。可以降到5-6,同时检查requests_cache是否生效——在session.get()后打印response.from_cache,如果是True说明命中缓存,重复请求会快很多。替换更轻量的解析工具
如果表格是静态渲染的,直接用requests+BeautifulSoup代替requests_html,速度会大幅提升,因为不需要启动无头浏览器。示例片段:import requests from bs4 import BeautifulSoup response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') table_tbody = soup.select_one("div:nth-of-type(5) table tbody") rows = table_tbody.find_all('tr')[:12] final_list = [row.get_text(strip=False).split("\n") for row in rows]清理冗余代码
原代码里的split_list函数定义了但没用到,循环里的i += 1也是多余的(range(2,14)已经自动递增),直接删掉即可。
内容的提问来源于stack exchange,提问作者Jacky

