如何使用Pandas将HTML表格标题作为MultiIndex纳入DataFrame?
问题:如何将HTML表格标题作为MultiIndex纳入Pandas DataFrame?
问题背景
我尝试将指定URL中的HTML表格读取到Pandas DataFrame中:
https://antismash-db.secondarymetabolites.org/output/GCF_006385935.1/
页面渲染后的表格包含N个我需要的表格,以及1个需排除的表格(标题以“No secondary metabolite”开头)。使用pd.read_html读取后得到3个表格,其中最后一个并非要排除的表格,而是我需要的、表头带有“NZ_”前缀的表格的拼接表。
我想知道:有没有办法将渲染表格的标题作为MultiIndex整合到DataFrame中?
手动实现示例
以下是手动添加MultiIndex的代码及效果:
# 读取HTML表格 dataframes = pd.read_html("https://antismash-db.secondarymetabolites.org/output/GCF_006385935.1/") # 将Region设为索引 dataframes = list(map(lambda df: df.set_index("Region"), dataframes)) # 手动添加标题和表格表头作为索引层级 dataframes[0].index = dataframes[0].index.map(lambda x: ("GCF_006385935.1", "NZ_CP041066.1", x)) dataframes[1].index = dataframes[1].index.map(lambda x: ("GCF_006385935.1", "NZ_CP041065.1", x)) # 合并表格 df_concat = pd.concat(dataframes[:-1], axis=0) # 将 替换为下划线 df_concat.index = df_concat.index.map(lambda x: (x[0], x[1], x[2].replace(" ","_"))) # 设置MultiIndex名称 df_concat.index.names = ["基因组编号", "序列编号", "区域"] df_concat
自动化解决方案
手动硬编码标题不够灵活,可通过BeautifulSoup抓取页面表格标题,自动对应到表格添加MultiIndex:
import pandas as pd import requests from bs4 import BeautifulSoup # 获取页面内容 url = "https://antismash-db.secondarymetabolites.org/output/GCF_006385935.1/" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # 筛选目标表格(排除标题以"No secondary metabolite"开头的) tables = [] for table in soup.find_all("table"): # 获取表格对应的前置标题标签 prev_tag = table.find_previous(["h3", "p"]) if prev_tag and not prev_tag.text.strip().startswith("No secondary metabolite"): seq_id = prev_tag.text.strip() # 解析表格为DataFrame df = pd.read_html(str(table))[0] tables.append((seq_id, df)) # 处理每个表格,添加MultiIndex processed_dfs = [] genome_id = "GCF_006385935.1" for seq_id, df in tables: df = df.set_index("Region") # 构建MultiIndex df.index = pd.MultiIndex.from_tuples( [(genome_id, seq_id, idx.replace(" ", "_")) for idx in df.index], names=["基因组编号", "序列编号", "区域"] ) processed_dfs.append(df) # 合并所有表格 final_df = pd.concat(processed_dfs, axis=0) print(final_df)
说明
- 通过
BeautifulSoup定位表格的前置标题,自动筛选出目标表格; - 无需手动硬编码标题,自动将标题作为MultiIndex的层级;
- 统一处理特殊字符替换,设置清晰的索引名称。
内容的提问来源于stack exchange,提问作者O.rka
相关产品推荐
相关产品推荐

