You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas使用drop_duplicates去重后导出CSV出现多余行列如何修复

问题原因
  • 新增行编号列:pandas的DataFrame.to_csv()方法默认会将DataFrame的行索引写入CSV作为第一列,未手动关闭该配置时就会额外多出一列行号。
  • 新增列编号首行:你读取文件时指定了header=None,说明原文件没有表头,pandas会自动为每列分配0、1、2……的默认列名,写入文件时未关闭表头输出,这些默认列名就会被当做首行写入。
修复方案

修改代码中写入CSV的行,添加参数关闭索引和表头输出即可:

def remove_duplicates(file):
    df = pd.read_csv(file, encoding="latin-1", header = None)
    Helper.printline(f"Rows in file {file}: {df.shape[0]}")
    df.drop_duplicates(keep='first', inplace=True)
    Helper.printline(f"Rows in file {file} with duplicates removed: {df.shape[0]}")
    # 仅修改下面这一行,添加两个参数
    df.to_csv(file, index=False, header=False)

内容的提问来源于stack exchange,提问作者arame3333

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 20:36:03