如何不依赖DataFrame索引使用pandas抓取指定动态HTML表格
问题原因
代码报错核心是两点:
pandas.read_html()的match参数默认仅匹配表格主体文本、<caption>标签内容、表格相邻的外部文本,不会扫描<thead>内的表头字段,因此传入"Comments"无法匹配到目标表格。- 目标表格没有固定索引,页面每日增减表格后索引会漂移,但表格外层存在固定的唯一标识,完全不需要靠索引或模糊文本匹配定位。
稳定定位方案
用BeautifulSoup先通过固定页面标识定位到目标表格元素,再将元素传入pandas.read_html()解析,完全不受页面其他表格增减、内容变动影响。
完整实现代码
import requests from bs4 import BeautifulSoup import pandas as pd url = "https://ciffc.net/en/ciffc/ext/member/sitrep/" # 加请求头模拟普通浏览器访问,避免被站点拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" } resp = requests.get(url, headers=headers) resp.encoding = resp.apparent_encoding soup = BeautifulSoup(resp.text, "lxml") # 定位逻辑:目标表格外层固定id为section-apl,内部表格容器id为apl_table_wrapper,是页面唯一标识,不会随每日内容更新变动 target_table_tag = soup.select_one("#section-apl #apl_table_wrapper table") # 将定位到的表格标签转为字符串传入pandas解析 apl_df = pd.read_html(str(target_table_tag))[0]
目标内容提取
不需要硬编码行号索引,直接按字段值筛选即可,不受表格行顺序变动影响:
# 提取育空地区(YT)的备注内容 yukon_comment = apl_df[apl_df["Agency"] == "YT"]["Comments"].iloc[0] print(yukon_comment)
运行后直接输出目标内容:
Yukon is at a level 3 prep level - but will trend upwards with the forecasted hot and dry weather.
备用定位逻辑
如果后续页面前端改版修改了外层id,可以通过表头特征遍历匹配表格,适配「表头包含Comments列」的识别特征:
all_tables = soup.find_all("table") target_table_tag = None for table in all_tables: th_text = [th.get_text(strip=True) for th in table.select("thead th")] # 同时匹配三个表头字段,避免和其他带Comments列的表格混淆 if {"Agency", "APL", "Comments"}.issubset(set(th_text)): target_table_tag = table break apl_df = pd.read_html(str(target_table_tag))[0]
如果静态requests请求拿到的源码中找不到对应表格,说明页面内容是前端JS动态渲染的,替换为可获取渲染后页面源码的请求方式即可,后续定位、解析逻辑完全不变。
内容的提问来源于stack exchange,提问作者gecco15
相关产品推荐
相关产品推荐

