You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pandas将HTM文件列表批量导出为xlsx格式文件?

报错原因
  • 核心问题:pd.read_html() 的返回值是DataFrame列表,会把HTM文件中识别到的所有表格按顺序存入列表,并不是单个DataFrame对象。列表类型没有to_excel()方法,因此会抛出AttributeError: 'list' object has no attribute 'to_excel'错误。
  • 原代码额外存在3个会导致运行失败的问题:
    • 用于生成文件名的count变量未定义
    • 文件名变量拼写不一致(先定义filesname,写入时用了filename)
    • 未校验输出目录是否存在,路径不存在时会触发写入报错
可直接运行的实现代码

以下代码适配「每个HTM文件内仅含1个需要提取的目标表格」的最常见场景:

import pandas as pd
import os
import glob

# 路径配置
input_path = r'D:/Phenology/02Climate/'
output_path = r'D:/Phenology/02Climate/Excel/'
# 自动创建输出目录,已存在时不报错
os.makedirs(output_path, exist_ok=True)

# 遍历目录下所有HTM文件
for htm_file in glob.glob(os.path.join(input_path, "*.htm")):
    # 读取HTM内所有表格
    table_list = pd.read_html(htm_file)
    # 提取第一个表格(如果目标表格不是第一个,修改索引值即可)
    df = table_list[0]
    
    # 生成对应xlsx文件名
    base_name = os.path.basename(htm_file).split('.')[0]
    save_file = os.path.join(output_path, f"{base_name}.xlsx")
    
    # 写入Excel,index=False用于取消写入pandas自动生成的行号列
    df.to_excel(save_file, index=False)
    print(f"转换完成:{save_file}")
特殊场景适配

如果单个HTM文件内包含多个需要保留的表格,可以把所有表格写入同一个Excel文件的不同工作表,替换循环内的写入逻辑即可:

with pd.ExcelWriter(save_file) as writer:
    for table_idx, table_df in enumerate(table_list):
        table_df.to_excel(writer, sheet_name=f"表格{table_idx+1}", index=False)

内容的提问来源于stack exchange,提问作者weraphong suaruang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 23:01:39