You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析多个相同id表格,提取含First goal stats的第6个表格

问题原因

你现有代码的问题在于,把findAll返回的所有btable表格列表整体转成了字符串,pd.read_html取[0]只会拿到所有表格中的第一个,自然匹配不到目标表格。
补充说明:HTML规范要求id属性全局唯一,但大量旧站点或未严格遵循规范的站点会重复使用相同id,你用匹配所有同id元素的逻辑是没问题的。

解决方案

方案1:按顺序取第6个表格

如果你确定目标表格固定是第6个,可以直接按索引取值(Python序列从0开始计数,第6个对应索引值为5):

site = requests.get(url, headers=headers)
soup = BeautifulSoup(site.content, 'html.parser')
# 获取所有id为btable的表格
tb_list = soup.find_all('table',{'id': 'btable'})
if len(tb_list) >= 6:
    # 取第6个表格
    target_tb = tb_list[5]
    df = pd.read_html(str(target_tb))[0]
else:
    print("页面中btable表格数量不足6个,请检查页面结构")

方案2:按文本内容匹配(更稳妥)

不需要依赖表格的排序规则,只要表格内包含指定文本就能匹配,避免页面结构调整后表格顺序变化导致的匹配失败:

site = requests.get(url, headers=headers)
soup = BeautifulSoup(site.content, 'html.parser')
tb_list = soup.find_all('table',{'id': 'btable'})
target_tb = None
for tb in tb_list:
    if "First goal stats:" in tb.get_text(strip=True):
        target_tb = tb
        break
if target_tb:
    df = pd.read_html(str(target_tb))[0]
else:
    print("未找到包含指定内容的目标表格")

内容的提问来源于stack exchange,提问作者Rod Mascarenhas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 05:57:02