You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas:重复表头行与CSV文件大小异常增大问题

问题解决方案

1. 重复表头行无法删除的问题

问题根源:fetch_data_from_api函数中,直接拼接多个API返回的CSV文本,每个CSV都自带表头。pd.read_csv仅会把第一行识别为表头,后续的表头行都会被当作普通数据行存入DataFrame。这些表头行的内容是字段名,和正常数据行内容差异明显,因此drop_duplicates()无法识别并删除它们。

修复方法有两种,任选其一即可:

方法一:手动处理文本拼接,跳过后续表头

def fetch_data_from_api(self):
    resource_dump_data = ""
    for idx, resource in enumerate(self.package["result"]["resources"]):
        if resource["datastore_active"]:
            url = self.base_url + "/datastore/dump/" + resource["id"]
            response_text = requests.get(url).text
            if idx == 0:
                # 第一个资源保留完整CSV(含表头)
                resource_dump_data += response_text
            else:
                # 后续资源跳过表头行,只取数据部分
                lines = response_text.splitlines()
                if len(lines) > 1:
                    resource_dump_data += "\n" + "\n".join(lines[1:])
    return pd.read_csv(io.StringIO(resource_dump_data))

方法二:单独读取每个资源的DataFrame再合并(更稳妥)

def fetch_data_from_api(self):
    dfs = []
    for resource in self.package["result"]["resources"]:
        if resource["datastore_active"]:
            url = self.base_url + "/datastore/dump/" + resource["id"]
            df = pd.read_csv(url)
            dfs.append(df)
    return pd.concat(dfs, ignore_index=True)

这种方式不需要手动处理文本,pandas会自动处理表头合并,从根源避免表头行混入数据。

2. 删除CSV后运行两次文件大小剧增的问题

问题根源:结合第一个问题来看,第一次运行时重复表头行被当作数据行存入CSV;第二次运行时,新拉取的数据又带入重复表头行,和已有数据concat后,drop_duplicates()无法去重(这些行和正常数据不重复),导致数据量叠加。另外,若API返回的多个资源本身存在重复数据,也会加剧这个问题。

修复步骤:

  1. 先解决上述重复表头问题,确保无表头行混入数据。
  2. 优化去重逻辑,可指定唯一标识字段提升准确性:
def load_data_to_main(self, new_data):
    if os.path.isfile(self.CSV_FILE_PATH):
        existing_data = pd.read_csv(self.CSV_FILE_PATH)
        combined = pd.concat([existing_data, new_data], ignore_index=True)
        # 如果数据有唯一标识字段(比如"id"),用subset参数指定,避免误删
        # combined = combined.drop_duplicates(subset=["id"], keep="last")
        combined = combined.drop_duplicates(keep="last")
    else:
        combined = new_data

    combined.to_csv(self.CSV_FILE_PATH, index=False)
  1. 对拉取的原始数据提前去重:
def fetch_data_from_api(self):
    dfs = []
    for resource in self.package["result"]["resources"]:
        if resource["datastore_active"]:
            url = self.base_url + "/datastore/dump/" + resource["id"]
            df = pd.read_csv(url)
            dfs.append(df)
    combined_df = pd.concat(dfs, ignore_index=True)
    # 对拉取的原始数据先去重
    combined_df = combined_df.drop_duplicates(keep="last")
    return combined_df

额外排查建议

写入CSV前可以打印关键信息,方便定位问题:

print(f"合并后总条数: {len(combined)}")
print(f"重复条数: {len(combined) - len(combined.drop_duplicates())}")

内容的提问来源于stack exchange,提问作者gitm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 16:42:42