如何在Python中按文件名升序合并同表头CSV文件并保留首文件表头?
我来帮你搞定这两个合并CSV的问题——文件顺序乱、列头截断,给你两种靠谱的解决方案:
解决方案1:使用Python标准库csv模块
这个方法不需要额外安装第三方库,用Python自带的工具就能完成,适合对环境有限制的场景:
import csv import os # 严格按照你要求的升序排列目标文件列表 target_files = [ 'AB201602.csv', 'AB201603.csv', 'AB201604.csv', 'AB201605.csv', 'AB201606.csv', 'AB201607.csv', 'AB201608.csv', 'AB201610.csv', 'AB201612.csv' ] output_filename = 'merged_result.csv' # 打开输出文件,指定编码和换行符避免格式问题 with open(output_filename, 'w', newline='', encoding='utf-8') as output_file: csv_writer = csv.writer(output_file, delimiter=',') is_first_file = True for file in target_files: # 检查文件是否存在,避免报错 if not os.path.exists(file): print(f"⚠️ 警告:文件 {file} 不存在,已自动跳过") continue with open(file, 'r', encoding='utf-8') as input_file: csv_reader = csv.reader(input_file, delimiter=',') if is_first_file: # 写入第一个文件的完整表头 header = next(csv_reader) csv_writer.writerow(header) is_first_file = False else: # 跳过后续文件的表头,只写数据行 next(csv_reader) # 逐行写入当前文件的数据 for row in csv_reader: csv_writer.writerow(row) print(f✅ 合并完成!输出文件:{output_filename}")
关键细节说明:
- 固定文件顺序:手动指定了有序的文件列表,彻底解决之前随机排序的问题,完全符合你要的升序要求。
- 避免列头截断:指定
encoding='utf-8'(如果你的CSV是GBK编码,改成encoding='gbk'),同时明确设置分隔符delimiter=','(如果是制表符分隔,改成'\t'),这两个参数是解决列头乱码/截断的核心。 - 兼容Windows格式:
newline=''参数可以避免Windows系统下合并后的文件出现多余空行。
解决方案2:使用pandas(更简洁高效)
如果你能安装第三方库,pandas会让合并操作更简单,代码量更少,还能自动处理列对齐:
import pandas as pd import os # 同样使用固定顺序的文件列表 target_files = [ 'AB201602.csv', 'AB201603.csv', 'AB201604.csv', 'AB201605.csv', 'AB201606.csv', 'AB201607.csv', 'AB201608.csv', 'AB201610.csv', 'AB201612.csv' ] output_filename = 'merged_result.csv' data_frames = [] for idx, file in enumerate(target_files): if not os.path.exists(file): print(f"⚠️ 警告:文件 {file} 不存在,已自动跳过") continue # 读取CSV文件,指定编码保证列头正常显示 df = pd.read_csv(file, encoding='utf-8') # 确保后续文件的列顺序和第一个文件完全一致(如果有列顺序变动的情况) if idx > 0: df = df[data_frames[0].columns] data_frames.append(df) # 合并所有数据,重置索引避免重复 merged_df = pd.concat(data_frames, ignore_index=True) # 保存结果,不写入索引列 merged_df.to_csv(output_filename, index=False, encoding='utf-8') print(f✅ 合并完成!输出文件:{output_filename}")
关键细节说明:
- 自动列对齐:pandas会自动匹配相同列名的列,就算后续文件列顺序和第一个不一样,也能正确合并。
- 更少的代码:不需要手动处理表头跳过和逐行写入,
pd.concat一步搞定合并。
你之前遇到问题的原因分析
- 文件顺序随机:之前可能用了
os.listdir()或glob.glob()获取文件,这些方法返回的顺序是系统默认的(不是按文件名升序),所以会出现乱序。手动指定有序列表是最稳妥的解决方式。 - 列头被截断:大概率是两个原因:
- 编码不匹配:比如文件实际是GBK编码,但你用UTF-8读取,导致字符乱码看起来像截断。
- 分隔符错误:比如CSV是制表符分隔,但你用逗号读取,导致多个列被合并成一个,视觉上像是列头被截断。
内容的提问来源于stack exchange,提问作者user9264558
相关产品推荐
相关产品推荐

