You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pd.insert()函数用法咨询:多CSV合并时前置标签列添加异常

问题分析与解决方案

你的代码主要有两个核心问题,导致label列没有按预期出现在输出CSV的开头:

  1. 循环中未累积处理后的DataFrame:你在循环里逐个处理每个CSV文件,但每次循环都会覆盖df变量,最后只保留了最后一个文件的处理结果,前面的文件数据全部丢失了。
  2. 写入时的结构转换错误:你用pd.DataFrame(dict(zip(df.index, df.value)), index=[0])转换数据的方式,会直接丢弃之前插入的label列,同时打乱原数据结构,导致label无法被正确写入输出文件。

修正后的完整代码

下面是调整后的代码,完美匹配你期望的输出格式:

import pandas as pd
import glob

path = r'.../Model_AMs'
allFiles = glob.glob(path + "/*.csv")
# 初始化列表存储所有处理好的DataFrame
processed_dfs = []

for file_ in allFiles:
    # 读取文件、转置、melt处理
    df = pd.read_csv(file_, header=None).astype(int).transpose().melt()
    # 根据文件名添加label列
    if 'slot' in file_:
        df.insert(loc=0, column='label', value=0)
        print('slot file!')
    else:
        df.insert(loc=0, column='label', value=1)
        print('boss file!')
    
    # 将长格式数据转成宽格式,得到label + pixel0/pixel1...的结构
    wide_df = df.pivot(columns='variable', values='value')
    # 把label列固定在第一位置
    wide_df.insert(0, 'label', df['label'].iloc[0])
    # 重置索引避免干扰
    wide_df = wide_df.reset_index(drop=True)
    # 加入结果列表
    processed_dfs.append(wide_df)

# 合并所有处理结果并写入CSV
final_df = pd.concat(processed_dfs, ignore_index=True)
final_df.to_csv('trainingdata.csv', index=False)

关键调整说明

  • 累积处理结果:用processed_dfs列表保存每个文件的处理结果,避免数据被循环覆盖。
  • 正确转换为宽格式:通过pivot()将melt后的长数据(每行对应一个像素值)转成宽数据(一行包含所有像素列),完美匹配你需要的输出结构。
  • 固定label列位置:转成宽格式后,手动将label列插入到第一列,确保它是输出CSV的首列。
  • 统一写入更高效:最后合并所有DataFrame后一次性写入,比循环追加更稳定,也更节省内存。

内存友好版(循环直接追加写入)

如果你的CSV文件总数据量很大,不想占用过多内存,可以在循环里直接处理并追加写入:

import pandas as pd
import glob

path = r'.../Model_AMs'
allFiles = glob.glob(path + "/*.csv")
# 标记是否为首次写入,控制表头输出
first_write = True

for file_ in allFiles:
    df = pd.read_csv(file_, header=None).astype(int).transpose().melt()
    label = 0 if 'slot' in file_ else 1
    df.insert(loc=0, column='label', value=label)
    print('slot file!' if label == 0 else 'boss file!')
    
    # 转宽格式并调整结构
    wide_df = df.pivot(columns='variable', values='value')
    wide_df.insert(0, 'label', label)
    wide_df = wide_df.reset_index(drop=True)
    
    # 追加写入,首次写入带表头,后续不重复写表头
    wide_df.to_csv('trainingdata.csv', index=False, mode='a', header=first_write)
    first_write = False

最终你会得到期望的输出格式:

label,pixel0,pixel1,pixel2,...
1,0,1,1,...
0,1,0,0,...
...

内容的提问来源于stack exchange,提问作者haagn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:41:12