Pandas读取合并多CSV时索引与格式异常问题求助
合并多个CSV文件时Pandas格式混乱问题解决
使用Pandas读取合并多个CSV文件时,出现额外索引导致最终格式混乱的问题,尝试过index_col=None、index_col=False、index_col=0等参数均无法解决。
原始CSV内容
- CSV1:
1704067200000,0.14720000,0.15300000,0 1704153600000,0.15200000,0.15600000,0
- CSV2:
1704758400000,0.13780000,0.13790000,0 1704844800000,0.13240000,0.13970000,0
期望合并结果
1704067200000,0.14720000,0.15300000,0 1704153600000,0.15200000,0.15600000,0 1704758400000,0.13780000,0.13790000,0 1704844800000,0.13240000,0.13970000,0
实际错误结果
,1704067200000,0.14720000,0.15300000,0,1704758400000,0.13780000,0.13790000 0,1704153600000.0,0.152,0.156,0,,, 1,,,,0,1704844800000.0,0.1324,0.1397
尝试过的代码
代码1
for folder in data_folders: data_lists = [pd.read_csv(csvfile, index_col=None) for csvfile in folder.glob('*.csv')] pd.concat(data_lists, ignore_index=True).to_csv(folder_1d / f"{coin_folder.name}.csv") print(data_lists)
打印的DataFrame结果:
[ 1704067200000 0.14720000 0.15300000 0 0 1704153600000 0.152 0.156 0, Unnamed: 0 1704067200000 0.14720000 0.15300000 0 1704758400000 \ 0 0 1.704154e+12 0.152 0.156 0 NaN 1 1 NaN NaN NaN 0 1.704845e+12 0.13780000 0.13790000 0 NaN NaN 1 0.1324 0.1397 , 1704758400000 0.13780000 0.13790000 0 0 1704844800000 0.1324 0.1397 0]
代码2
all_files = glob.glob(os.path.join(folder_1d, "*.csv")) df = pd.concat((pd.read_csv(f, index_col=False) for f in all_files), ignore_index=True)
得到的结果:
1704067200000 0.14720000 0.15300000 0 1704758400000 0.13780000 \ 0 1.704154e+12 0.152 0.156 0 NaN NaN 1 NaN NaN NaN 0 1.704845e+12 0.1324 0.13790000 0 NaN 1 0.1397
解决方案
问题根源有两个:
- 部分CSV文件被错误写入了索引列(打印结果中的
Unnamed: 0列),说明之前保存CSV时未设置index=False,导致文件多了一列索引。 - 原始CSV无表头,Pandas默认将第一行数据当作列名,导致不同CSV的列名不匹配,合并时出现错位和NaN。
修复步骤
1. 读取CSV时指定无表头
读取时设置header=None,让Pandas用统一的数字列名,保证所有CSV的列结构一致:
pd.read_csv(csvfile, header=None, index_col=None)
2. 保存CSV时禁止写入索引和表头
使用to_csv时必须加index=False避免生成额外索引,加header=False避免输出表头(因为原始CSV无表头):
merged_df.to_csv(output_path, index=False, header=False)
完整修复代码
for folder in data_folders: # 读取所有CSV,指定无表头、不使用索引列 data_lists = [pd.read_csv(csvfile, header=None, index_col=None) for csvfile in folder.glob('*.csv')] # 合并DataFrame merged_df = pd.concat(data_lists, ignore_index=True) # 保存合并结果,禁止写入索引和表头 merged_df.to_csv(folder_1d / f"{coin_folder.name}.csv", index=False, header=False)
额外处理
如果已有带额外索引列的旧CSV文件,可读取时跳过第一列:
pd.read_csv(csvfile, header=None, index_col=None, usecols=[1,2,3,4])
更彻底的方式是重新生成不带索引的原始CSV文件。
内容的提问来源于stack exchange,提问作者masterpiece
相关产品推荐
相关产品推荐

