如何用Python抓取含Load More选项的网页完整表格?
抓取带Load More的完整表格解决方案
问题背景
需要抓取https://www.mykhel.com/football/indian-super-league-player-stats-l750/上的完整球员统计表格,但使用requests和pandas只能获取默认加载的第一页数据,无法获取Load More按钮加载的后续内容,表格末尾显示“Load More....”。
原代码:
import pandas as pd import requests from six.moves import urllib URL2 = "https://www.mykhel.com/football/indian-super-league-player-stats-l750/" header = {'Accept-Language': "en-US,en;q=0.9", 'User-Agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 " "(KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36" } resp2 = requests.get(url=URL2, headers=header).text tables2 = pd.read_html(resp2) overview_table2= tables2[0] overview_table2
解决思路
这个网站的Load More按钮是通过AJAX POST请求加载后续数据的,我们可以直接模拟这些请求来获取所有分页数据,无需使用浏览器自动化工具。
步骤1:分析请求参数
打开浏览器开发者工具(F12),点击Load More按钮,查看Network面板的XHR请求,发现核心接口是https://www.mykhel.com/matchCenter/statsLoadMore.php,POST请求的关键参数:
page:分页页码(从2开始,第一页是默认加载的)type:固定为playermatchId:赛事ID,这里是750statsType:统计类型,这里是overviewgameFormat:固定为league
步骤2:完整爬取代码
import pandas as pd import requests # 模拟浏览器请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36', 'Accept-Language': 'en-US,en;q=0.9', 'X-Requested-With': 'XMLHttpRequest' # 标记为AJAX请求,避免被拦截 } # 1. 获取第一页默认数据 base_url = "https://www.mykhel.com/football/indian-super-league-player-stats-l750/" resp = requests.get(base_url, headers=headers).text first_page_tables = pd.read_html(resp) full_stats_df = first_page_tables[0].iloc[:-1] # 移除末尾的"Load More...."行 # 2. 循环加载后续分页数据 api_url = "https://www.mykhel.com/matchCenter/statsLoadMore.php" current_page = 2 while True: payload = { 'page': current_page, 'type': 'player', 'matchId': '750', 'statsType': 'overview', 'gameFormat': 'league', 'seasonId': '' } try: # 发送POST请求获取分页数据 api_resp = requests.post(api_url, headers=headers, data=payload) api_resp.raise_for_status() # 检查请求是否成功 # 解析返回的HTML表格 page_tables = pd.read_html(api_resp.text) if not page_tables: break # 没有更多数据时退出循环 page_df = page_tables[0] # 检查是否是最后一页,移除"Load More"行 if "Load More" in page_df.iloc[-1].values: page_df = page_df.iloc[:-1] # 合并到总数据框 full_stats_df = pd.concat([full_stats_df, page_df], ignore_index=True) current_page += 1 except requests.exceptions.RequestException as e: print(f"请求出错,停止加载:{e}") break # 输出完整数据 print(full_stats_df) # 保存为CSV文件 full_stats_df.to_csv("indian_super_league_full_player_stats.csv", index=False)
代码说明
- 先获取第一页数据并清理掉末尾的Load More行
- 循环发送POST请求,每次递增页码,直到没有更多数据返回
- 自动处理最后一页的Load More行,确保数据完整
- 最终将所有数据合并为一个DataFrame,可直接保存为CSV
内容的提问来源于stack exchange,提问作者Footilytics - Indian football
相关产品推荐
相关产品推荐

