如何将循环生成的数组保存为CSV多列而非追加为单列?
解决循环中给CSV添加新列的问题
嘿,我明白你现在的困扰了——每次循环生成的单列数组,要么被追加成一堆行,要么覆盖掉之前的列,完全达不到添加新列的效果对吧?
先给你捋清楚问题出在哪:你当前用mode='a'写入CSV的时候,Pandas是在文件末尾追加行数据,而不是添加新的列。如果去掉mode='a',每次写入都会直接覆盖整个文件,旧列自然就没了,这肯定不是你想要的结果。
下面给你两种解决方案,按需选择:
方案一:先收集所有列,最后一次性保存(推荐)
这种方法效率最高,也最不容易出错,适合数据量在内存能承受的情况:
import pandas as pd # 初始化空DataFrame,用来存所有迭代生成的列 predict_df = pd.DataFrame() ground_truth_df = pd.DataFrame() for ij in range(0, 100): # 这里替换成你生成predict_result和ground_truth的代码 # predict_result = 你的生成逻辑(shape [100,1]) # ground_truth = 你的生成逻辑(shape [100,1]) # 把二维数组转成一维,添加为新列,列名可以用迭代次数区分 predict_df[f'iter_{ij}'] = predict_result.flatten() ground_truth_df[f'iter_{ij}'] = ground_truth.flatten() # 最后统一保存,index=False是去掉Pandas默认的行索引 predict_df.to_csv('predict_results.csv', index=False) ground_truth_df.to_csv('ground_truths.csv', index=False)
核心思路就是先把所有要加的列都存在内存里的DataFrame中,最后一次性写入文件,避免反复读写带来的问题。
方案二:循环中逐步读写文件(适合超大数据)
如果你的数据量特别大,没法一次性存在内存里,那可以每次循环读取已有CSV,合并新列后再保存:
import pandas as pd for ij in range(0, 100): # 生成你的predict_result和ground_truth # predict_result = ... # ground_truth = ... # 处理预测结果文件 try: # 文件已存在,读取现有数据 existing_predict = pd.read_csv('predict_results.csv') # 添加新列 existing_predict[f'iter_{ij}'] = predict_result.flatten() existing_predict.to_csv('predict_results.csv', index=False) except FileNotFoundError: # 第一次循环文件不存在,直接保存第一列 pd.DataFrame(predict_result.flatten(), columns=[f'iter_{ij}']).to_csv('predict_results.csv', index=False) # 处理真实标签文件,逻辑和上面完全一致 try: existing_truth = pd.read_csv('ground_truths.csv') existing_truth[f'iter_{ij}'] = ground_truth.flatten() existing_truth.to_csv('ground_truths.csv', index=False) except FileNotFoundError: pd.DataFrame(ground_truth.flatten(), columns=[f'iter_{ij}']).to_csv('ground_truths.csv', index=False)
这种方法每次循环都要读写文件,效率会低一些,但胜在内存占用小,适合超大数据集。
内容的提问来源于stack exchange,提问作者user288609
相关产品推荐
相关产品推荐

