You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python抓取含Load More选项的网页完整表格?

抓取带Load More的完整表格解决方案

问题背景

需要抓取https://www.mykhel.com/football/indian-super-league-player-stats-l750/上的完整球员统计表格,但使用requests和pandas只能获取默认加载的第一页数据,无法获取Load More按钮加载的后续内容,表格末尾显示“Load More....”。

原代码:

import pandas as pd
import requests
from six.moves import urllib

URL2 = "https://www.mykhel.com/football/indian-super-league-player-stats-l750/"
header = {'Accept-Language': "en-US,en;q=0.9",
          'User-Agent': "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 "
                        "(KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36"
          }

resp2 = requests.get(url=URL2, headers=header).text

tables2 = pd.read_html(resp2)
overview_table2= tables2[0]
overview_table2

解决思路

这个网站的Load More按钮是通过AJAX POST请求加载后续数据的,我们可以直接模拟这些请求来获取所有分页数据,无需使用浏览器自动化工具。

步骤1:分析请求参数

打开浏览器开发者工具(F12),点击Load More按钮,查看Network面板的XHR请求,发现核心接口是https://www.mykhel.com/matchCenter/statsLoadMore.php,POST请求的关键参数:

  • page:分页页码(从2开始,第一页是默认加载的)
  • type:固定为player
  • matchId:赛事ID,这里是750
  • statsType:统计类型,这里是overview
  • gameFormat:固定为league

步骤2:完整爬取代码

import pandas as pd
import requests

# 模拟浏览器请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/107.0.0.0 Safari/537.36',
    'Accept-Language': 'en-US,en;q=0.9',
    'X-Requested-With': 'XMLHttpRequest'  # 标记为AJAX请求,避免被拦截
}

# 1. 获取第一页默认数据
base_url = "https://www.mykhel.com/football/indian-super-league-player-stats-l750/"
resp = requests.get(base_url, headers=headers).text
first_page_tables = pd.read_html(resp)
full_stats_df = first_page_tables[0].iloc[:-1]  # 移除末尾的"Load More...."行

# 2. 循环加载后续分页数据
api_url = "https://www.mykhel.com/matchCenter/statsLoadMore.php"
current_page = 2

while True:
    payload = {
        'page': current_page,
        'type': 'player',
        'matchId': '750',
        'statsType': 'overview',
        'gameFormat': 'league',
        'seasonId': ''
    }

    try:
        # 发送POST请求获取分页数据
        api_resp = requests.post(api_url, headers=headers, data=payload)
        api_resp.raise_for_status()  # 检查请求是否成功

        # 解析返回的HTML表格
        page_tables = pd.read_html(api_resp.text)
        if not page_tables:
            break  # 没有更多数据时退出循环

        page_df = page_tables[0]
        # 检查是否是最后一页,移除"Load More"行
        if "Load More" in page_df.iloc[-1].values:
            page_df = page_df.iloc[:-1]

        # 合并到总数据框
        full_stats_df = pd.concat([full_stats_df, page_df], ignore_index=True)
        current_page += 1

    except requests.exceptions.RequestException as e:
        print(f"请求出错,停止加载:{e}")
        break

# 输出完整数据
print(full_stats_df)
# 保存为CSV文件
full_stats_df.to_csv("indian_super_league_full_player_stats.csv", index=False)

代码说明

  • 先获取第一页数据并清理掉末尾的Load More行
  • 循环发送POST请求,每次递增页码,直到没有更多数据返回
  • 自动处理最后一页的Load More行,确保数据完整
  • 最终将所有数据合并为一个DataFrame,可直接保存为CSV

内容的提问来源于stack exchange,提问作者Footilytics - Indian football

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 17:30:42