Python循环向列表追加元素异常:重复存储最后lot_id数据
问题分析:循环中Numpy数组添加到列表后全部被最后一次值覆盖
这个问题的核心是Numpy数组的引用特性在搞鬼,我来给你一步步拆解:
你代码里的temp是一个预先定义好的Numpy数组对象,在循环过程中,你只是不断修改这个数组的内容,然后把它添加到lsttest2列表中。但Numpy数组属于可变对象,列表存储的是对象的引用(也就是内存地址),而不是对象的副本。
换句话说,你每次循环都是把同一个temp的"地址"加到列表里,后续修改temp的内容时,列表里所有之前添加的"元素"其实都是指向这个同一个数组的。等循环结束,所有引用自然都指向最后一次修改后的G的填充数据,就出现了你看到的异常。
两种可行的解决方案
方案1:每次循环创建全新的零数组
把temp的定义移到循环内部,这样每次循环都会生成一个独立的零数组,修改后添加到列表的是完全不同的对象:
lotids = dummy['lot_id'].unique() lsttest2 = [] #define empty list lots = len(lotids) for i in range(lots): lot = lotids[i] lotsub = dummy[dummy.lot_id==lot] lotsub = lotsub.to_numpy() lotsub = np.delete(lotsub, 0, 1) #remove the lot_id col temp = np.zeros(shape=(10,4)) # 每次循环新建一个零数组 temp[:lotsub.shape[0],:lotsub.shape[1]] = lotsub #add it to the all zero array to pad out print(temp.shape) print(temp) lsttest2.append(temp) print("this is the end of loop---", i, " the list now has ", len(lsttest2), " elements") print(lsttest2)
方案2:添加数组的副本到列表
如果想复用外部的模板结构,那每次添加的时候要创建当前temp的副本,这样列表里存储的是独立的数组对象,后续修改原temp不会影响已添加的元素:
lotids = dummy['lot_id'].unique() temp = np.zeros(shape=(10,4)) #template to pad to lsttest2 = [] #define empty list lots = len(lotids) for i in range(lots): lot = lotids[i] lotsub = dummy[dummy.lot_id==lot] lotsub = lotsub.to_numpy() lotsub = np.delete(lotsub, 0, 1) #remove the lot_id col temp[:lotsub.shape[0],:lotsub.shape[1]] = lotsub #add it to the all zero array to pad out print(temp.shape) print(temp) lsttest2.append(temp.copy()) # 添加副本而不是原引用 print("this is the end of loop---", i, " the list now has ", len(lsttest2), " elements") print(lsttest2)
两种方案都能让列表里的每个元素对应各自lot_id的填充数据,不会被最后一次循环的内容覆盖。
内容的提问来源于stack exchange,提问作者PaulBeales
相关产品推荐
相关产品推荐

