You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将HTML表格中'Total Electric Industry'部分导入Pandas DataFrame?

问题解决:定位目标表格区域失败的修复方案

问题原因

直接用精确字符串匹配"Total Electric Industry"和"Full-Service Providers"无法找到对应行,是因为网页表格中的目标文本可能带有不可见空白字符(如前后空格、换行符),或是跨列合并单元格在pandas读取后存在格式差异,导致精确匹配失效。

修复步骤与代码

1. 先排查表格实际内容

运行以下代码查看表格前20行,确认目标文本的位置和格式:

# Total residential customers
us_res_html = 'https://www.eia.gov/electricity/annual/html/epa_02_01.html'
us_res = pd.read_html(us_res_html)
us_res = us_res[1]

# 打印前20行排查内容
print(us_res.head(20))

2. 改用模糊匹配定位索引

替换原索引查找代码,使用字符串包含匹配(忽略空白和格式差异):

# 查找目标行索引(模糊匹配)
start_idx = us_res[us_res[us_res.columns[0]].str.contains('Total Electric Industry', na=False)].index
end_idx = us_res[us_res[us_res.columns[0]].str.contains('Full-Service Providers', na=False)].index

# 若单列匹配失败,尝试遍历所有列查找
def find_target_row(df, target_text):
    for col in df.columns:
        match_rows = df[df[col].str.contains(target_text, na=False)]
        if not match_rows.empty:
            return match_rows.index[0]
    return None

# 调用函数获取索引
start_idx = find_target_row(us_res, 'Total Electric Industry')
end_idx = find_target_row(us_res, 'Full-Service Providers')

3. 提取目标表格区域

确认索引有效后,提取并整理目标数据:

if start_idx is not None and end_idx is not None:
    # 提取从start到end前一行的内容
    total_electric_df = us_res.loc[start_idx:end_idx-1]
    # 重置索引
    total_electric_df = total_electric_df.reset_index(drop=True)
    print("成功提取目标表格:")
    print(total_electric_df)
else:
    print("未找到目标行,请检查表格内容")

关键说明

  • str.contains() 会匹配包含目标文本的所有行,避免了精确匹配对空白字符的敏感问题
  • 遍历列的函数可以处理跨列合并单元格导致文本不在第一列的情况

内容的提问来源于stack exchange,提问作者Mainland

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 22:40:55