为何for循环在处理大型数据集时无法正常工作?
问题场景
我在10条记录的小型数据集上测试代码,数据集如下:
预期的LOST PACKETS属性处理效果为:
异常代码
for lost_p, plr in zip(df['LOST PACKETS'], df['PACKET LOSS RATIO']): if (lost_p == 0): df['LOST PACKETS'] = df['LOST PACKETS'].replace(lost_p, 0) elif (lost_p > 0 and plr > 0.07): df['LOST PACKETS'] = df['LOST PACKETS'].replace(lost_p, 1) elif (lost_p > 0 and plr < 0.07): df['LOST PACKETS'] = df['LOST PACKETS'].replace(lost_p, 0) else: df['LOST PACKETS'] = df['LOST PACKETS'].replace(lost_p, 1)
实际问题
我处理的数据集:
代码执行后,LOST PACKETS列本该显示1的位置全是0:
执行df['LOST PACKETS'].value_counts()结果为:
0 14760 Name: LOST PACKETS, dtype: int64
问题原因
这段代码的核心错误是用循环遍历每行后,调用replace修改整个列的所有匹配值:
- 当某一行
lost_p>0且plr>0.07时,会把整个列里所有等于该lost_p的值替换成1; - 但如果后续有另一行相同的
lost_p但plr<0.07,又会把整个列里所有该lost_p的值替换成0; - 最终所有值都会被后续的替换操作覆盖,导致全变成0。
另外,lost_p ==0时替换0完全是多余操作。
正确写法
Pandas处理行级条件判断,应该用向量化操作,避免循环,既高效又不会出错:
方法1:使用numpy.where
import numpy as np # 定义条件 cond1 = df['LOST PACKETS'] == 0 cond2 = (df['LOST PACKETS'] > 0) & (df['PACKET LOSS RATIO'] > 0.07) cond3 = (df['LOST PACKETS'] > 0) & (df['PACKET LOSS RATIO'] < 0.07) # 按条件赋值 df['LOST PACKETS'] = np.where(cond1, 0, np.where(cond2, 1, np.where(cond3, 0, 1)))
方法2:使用apply(适合复杂逻辑)
def update_lost_packets(row): if row['LOST PACKETS'] == 0: return 0 elif row['PACKET LOSS RATIO'] > 0.07: return 1 elif row['PACKET LOSS RATIO'] < 0.07: return 0 else: return 1 df['LOST PACKETS'] = df.apply(update_lost_packets, axis=1)
方法3:链式条件赋值
# 先初始化默认值1 df['LOST PACKETS'] = 1 # 满足条件设为0 df.loc[df['LOST PACKETS'] == 0, 'LOST PACKETS'] = 0 df.loc[(df['LOST PACKETS'] > 0) & (df['PACKET LOSS RATIO'] < 0.07), 'LOST PACKETS'] = 0
这些方法都是针对每行单独判断赋值,不会出现全局替换覆盖的问题。
内容的提问来源于stack exchange,提问作者K___V
相关产品推荐
相关产品推荐

