使用numpy insert时未插入行反而替换行的问题排查
问题:numpy insert未实现行插入,反而出现替换效果,行数未增加
尝试编写向CSV文件插入行的程序,使用numpy.insert实现,但实际操作是替换行而非插入,最终输出的CSV行数未增加(预期11行,实际仅10行)。
当前代码
import numpy as np import pandas as pd #import the csv file file = np.loadtxt("name.csv", skiprows=1, dtype='<U70', delimiter =',') #get row and col count fileShape = file.shape rows = fileShape[0] cols = fileShape[1] #iterate over array for row in range(rows): for col in range(cols): #the only column that matters to me is the 5th one of my csv #also blocking overflow with the row check if (col == 4 and row + 1 < rows): #if the current row does not equal the next row #then I want to insert a row to do some things in Excel with if (link[row][col] != link[row+1][col]): #grab the next row and store it temp = file[row+1] #replace the 6th column with nothing temp[5] = "" #insert the new row into the array np.insert(file, row, [temp], 0) outfile = pd.DataFrame(file) outfile.to_csv("OutFile.csv")
输入CSV内容
ccType,number,date,payee,total,indAmt,memo,category mastercard,30,11/21/2022,Bluejam,287.24,44.33,,Sports mastercard,30,11/23/2022,Fanoodle,287.24,95.95,,Health mastercard,30,11/25/2022,Eazzy,287.24,1.2,,Automotive mastercard,30,11/26/2022,Dabfeed,287.24,68.97,,Games mastercard,30,11/30/2022,Jaloo,287.24,76.79,,Games mastercard,50,7/4/2023,Shufflebeat,317.13,91.91,,Sports mastercard,50,7/4/2023,Meembee,317.13,94.69,,Toys mastercard,50,7/5/2023,Jabberbean,317.13,67.01,,Computers mastercard,50,7/28/2023,Wikibox,317.13,33.18,,Movies mastercard,50,7/29/2023,Shufflebeat,317.13,30.34,,Automotive
实际输出CSV内容
ccType,number,date,payee,total,indAmt,memo,category mastercard,30,11/21/2022,Bluejam,287.24,44.33,,Sports mastercard,30,11/23/2022,Fanoodle,287.24,95.95,,Health mastercard,30,11/25/2022,Eazzy,287.24,1.2,,Automotive mastercard,30,11/26/2022,Dabfeed,287.24,68.97,,Games mastercard,30,11/30/2022,Jaloo,287.24,76.79,,Games mastercard,50,7/4/2023,Shufflebeat,317.13,,,Sports mastercard,50,7/4/2023,Meembee,317.13,94.69,,Toys mastercard,50,7/5/2023,Jabberbean,317.13,67.01,,Computers mastercard,50,7/28/2023,Wikibox,317.13,33.18,,Movies mastercard,50,7/29/2023,Shufflebeat,317.13,30.34,,Automotive
注意:输出中mastercard,50,7/4/2023,Shufflebeat,317.13,,,Sports是替换了原行,而非插入,行数未增加。
问题原因与解决方案
核心问题点
- numpy.insert不修改原数组:
np.insert()会返回一个插入后的新数组,不会直接修改原数组file,必须将返回值重新赋值给file才会生效。 - 变量名错误:代码中使用了未定义的
link变量,应替换为实际的数组变量file。 - 插入位置错误:要在当前行和下一行之间插入新行,插入位置应为
row+1,而非row。 - 循环遍历逻辑问题:正向遍历原数组行数时,插入行后数组长度增加会导致后续索引错位,需改为反向遍历。
修正后的代码
import numpy as np import pandas as pd # 单独读取表头,避免插入行时丢失表头 header = np.loadtxt("name.csv", max_rows=1, dtype='<U70', delimiter=',') file = np.loadtxt("name.csv", skiprows=1, dtype='<U70', delimiter=',') rows = file.shape[0] # 反向遍历行索引,防止插入新行后原索引失效 for row in range(rows-2, -1, -1): # 检查第5列(索引4)当前行与下一行是否不同 if file[row][4] != file[row+1][4]: # 复制下一行并清空第6列(索引5) temp = file[row+1].copy() temp[5] = "" # 插入新行到row+1位置,并更新原数组 file = np.insert(file, row+1, temp, axis=0) # 合并表头和处理后的数据 result = np.vstack([header, file]) # 保存为CSV,禁用索引和默认表头 outfile = pd.DataFrame(result) outfile.to_csv("OutFile.csv", index=False, header=False)
修正说明
- 单独读取表头,确保输出CSV保留原始表头结构
- 反向遍历避免插入行后索引错位问题
- 将
np.insert()的返回值赋值给file,确保原数组被更新 - 使用
copy()避免修改原数组中的行数据 - 最终合并表头与数据,保证输出格式正确
内容的提问来源于stack exchange,提问作者C0deGyver
相关产品推荐
相关产品推荐

