You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何不依赖DataFrame索引使用pandas抓取指定动态HTML表格

问题原因

代码报错核心是两点:

  • pandas.read_html()的match参数默认仅匹配表格主体文本、<caption>标签内容、表格相邻的外部文本,不会扫描<thead>内的表头字段,因此传入"Comments"无法匹配到目标表格。
  • 目标表格没有固定索引,页面每日增减表格后索引会漂移,但表格外层存在固定的唯一标识,完全不需要靠索引或模糊文本匹配定位。
稳定定位方案

用BeautifulSoup先通过固定页面标识定位到目标表格元素,再将元素传入pandas.read_html()解析,完全不受页面其他表格增减、内容变动影响。

完整实现代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

url = "https://ciffc.net/en/ciffc/ext/member/sitrep/"
# 加请求头模拟普通浏览器访问,避免被站点拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
}

resp = requests.get(url, headers=headers)
resp.encoding = resp.apparent_encoding
soup = BeautifulSoup(resp.text, "lxml")

# 定位逻辑:目标表格外层固定id为section-apl,内部表格容器id为apl_table_wrapper,是页面唯一标识,不会随每日内容更新变动
target_table_tag = soup.select_one("#section-apl #apl_table_wrapper table")
# 将定位到的表格标签转为字符串传入pandas解析
apl_df = pd.read_html(str(target_table_tag))[0]

目标内容提取

不需要硬编码行号索引,直接按字段值筛选即可,不受表格行顺序变动影响:

# 提取育空地区(YT)的备注内容
yukon_comment = apl_df[apl_df["Agency"] == "YT"]["Comments"].iloc[0]
print(yukon_comment)

运行后直接输出目标内容:

Yukon is at a level 3 prep level - but will trend upwards with the forecasted hot and dry weather.

备用定位逻辑

如果后续页面前端改版修改了外层id,可以通过表头特征遍历匹配表格,适配「表头包含Comments列」的识别特征:

all_tables = soup.find_all("table")
target_table_tag = None
for table in all_tables:
    th_text = [th.get_text(strip=True) for th in table.select("thead th")]
    # 同时匹配三个表头字段,避免和其他带Comments列的表格混淆
    if {"Agency", "APL", "Comments"}.issubset(set(th_text)):
        target_table_tag = table
        break

apl_df = pd.read_html(str(target_table_tag))[0]

如果静态requests请求拿到的源码中找不到对应表格,说明页面内容是前端JS动态渲染的,替换为可获取渲染后页面源码的请求方式即可,后续定位、解析逻辑完全不变。

内容的提问来源于stack exchange,提问作者gecco15

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 04:42:10