如何批量将扩展名是.xls的HTML文件转换为Excel XLSX文件
批量转换伪装为.xls格式的HTML文件到Excel
你的现有代码只能处理单个文件,要实现批量转换可以参考以下两种方案:
方案一:自动遍历指定目录下的所有目标文件
这种方式适合处理某个文件夹里所有伪装成.xls的HTML文件,无需手动逐个输入路径:
import pandas as pd import os # 配置源目录(存放待转换文件的文件夹)和输出目录(保存转换后文件的文件夹) source_directory = "/Users/abhijitchavan/Desktop/yes" output_directory = "/Users/abhijitchavan/Desktop/yes/converted_files" # 确保输出目录存在,不存在则自动创建 os.makedirs(output_directory, exist_ok=True) # 遍历源目录下的所有文件 for filename in os.listdir(source_directory): # 只处理.xls后缀的文件 if filename.endswith(".xls"): full_file_path = os.path.join(source_directory, filename) # 读取HTML格式的表格(取第一个表格) tables = pd.read_html(full_file_path) target_table = tables[0] # 生成输出文件名:原文件名替换后缀为.xlsx output_filename = os.path.splitext(filename)[0] + ".xlsx" full_output_path = os.path.join(output_directory, output_filename) # 保存为Excel文件,去掉默认的索引列 target_table.to_excel(full_output_path, index=False) print(f"完成转换:{full_output_path}")
方案二:手动指定多个文件路径(适合零散文件)
如果需要处理的文件分散在不同位置,可以通过命令行传入多个文件路径:
import pandas as pd import sys import os # 检查是否传入了文件路径 if len(sys.argv) < 2: print("使用方式:python 你的脚本名.py 文件1.xls 文件2.xls 文件3.xls") sys.exit(1) # 遍历所有传入的文件路径 for file_path in sys.argv[1:]: # 检查文件是否存在 if not os.path.exists(file_path): print(f"跳过:文件不存在 - {file_path}") continue # 只处理.xls后缀的文件 if not file_path.endswith(".xls"): print(f"跳过:非.xls格式文件 - {file_path}") continue # 读取并转换 tables = pd.read_html(file_path) target_table = tables[0] output_path = os.path.splitext(file_path)[0] + ".xlsx" target_table.to_excel(output_path, index=False) print(f"完成转换:{output_path}")
使用说明:
- 把代码保存成
.py文件(比如batch_convert.py) - 在命令行运行:
python batch_convert.py /path/to/file1.xls /path/to/file2.xls
注意点:
- 确保所有
.xls文件实际是HTML格式,否则pd.read_html会报错 - 如果文件内包含多个表格,代码默认取第一个,若需调整可修改
tables[0]的索引值 - 方案一的输出目录可以自定义,避免覆盖原文件
内容的提问来源于stack exchange,提问作者abhijit chavan
相关产品推荐
相关产品推荐

