Python Pandas使用drop_duplicates去重后导出CSV出现多余行列如何修复
问题原因
- 新增行编号列:
pandas的DataFrame.to_csv()方法默认会将DataFrame的行索引写入CSV作为第一列,未手动关闭该配置时就会额外多出一列行号。 - 新增列编号首行:你读取文件时指定了
header=None,说明原文件没有表头,pandas会自动为每列分配0、1、2……的默认列名,写入文件时未关闭表头输出,这些默认列名就会被当做首行写入。
修复方案
修改代码中写入CSV的行,添加参数关闭索引和表头输出即可:
def remove_duplicates(file): df = pd.read_csv(file, encoding="latin-1", header = None) Helper.printline(f"Rows in file {file}: {df.shape[0]}") df.drop_duplicates(keep='first', inplace=True) Helper.printline(f"Rows in file {file} with duplicates removed: {df.shape[0]}") # 仅修改下面这一行,添加两个参数 df.to_csv(file, index=False, header=False)
内容的提问来源于stack exchange,提问作者arame3333
相关产品推荐
相关产品推荐

