使用BeautifulSoup提取指定h1下的表格并存储为Pandas DataFrame
解决方案
核心调整思路是直接定位到文本为Tables的h1节点,仅筛选该节点之后出现的表格,自动过滤前置无效内容,修改后的完整代码如下:
import pandas as pd from bs4 import BeautifulSoup soup = BeautifulSoup(self.body, features="lxml") # 定位目标<h1>Tables</h1>节点 target_h1 = soup.find('h1', string='Tables') if not target_h1: print("Page doesn't contain tables") else: # 提取目标h1之后的所有h2作为表格标题 table_headers = [tag.text for tag in target_h1.find_all_next('h2')] # 仅提取目标h1之后的所有表格,排除前置无效表格 tables_raw = [[[cell.text for cell in row("th") + row("td")] for row in table("tr")] for table in target_h1.find_all_next('table')] # 生成DataFrame并关联标题,逻辑保持不变 tables_df = [pd.DataFrame(table) for table in tables_raw] tables_and_names = list(zip(table_headers, tables_df))
调整说明
- 不再先提取全页所有h1/h2标题再做索引切片,改为直接定位目标h1节点,避免全页无关标题干扰
- 利用BeautifulSoup内置的
find_all_next()方法,自动获取当前节点之后的所有符合条件的元素,天然过滤了h1之前的无效表格,同时保证了h2标题和后续表格的顺序一一对应,不会出现索引错位问题
内容的提问来源于stack exchange,提问作者Omega
相关产品推荐
相关产品推荐

