You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup提取指定h1下的表格并存储为Pandas DataFrame

解决方案

核心调整思路是直接定位到文本为Tables的h1节点,仅筛选该节点之后出现的表格,自动过滤前置无效内容,修改后的完整代码如下:

import pandas as pd
from bs4 import BeautifulSoup

soup = BeautifulSoup(self.body, features="lxml")
# 定位目标<h1>Tables</h1>节点
target_h1 = soup.find('h1', string='Tables')

if not target_h1:
    print("Page doesn't contain tables")
else:
    # 提取目标h1之后的所有h2作为表格标题
    table_headers = [tag.text for tag in target_h1.find_all_next('h2')]
    # 仅提取目标h1之后的所有表格,排除前置无效表格
    tables_raw = [[[cell.text for cell in row("th") + row("td")] for row in table("tr")] for table in target_h1.find_all_next('table')]
    # 生成DataFrame并关联标题,逻辑保持不变
    tables_df = [pd.DataFrame(table) for table in tables_raw]
    tables_and_names = list(zip(table_headers, tables_df))

调整说明

  • 不再先提取全页所有h1/h2标题再做索引切片,改为直接定位目标h1节点,避免全页无关标题干扰
  • 利用BeautifulSoup内置的find_all_next()方法,自动获取当前节点之后的所有符合条件的元素,天然过滤了h1之前的无效表格,同时保证了h2标题和后续表格的顺序一一对应,不会出现索引错位问题

内容的提问来源于stack exchange,提问作者Omega

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 02:24:03