为何用iterrows()行对象赋值首次有效后续失效?
问题:使用
iterrows()返回的行对象赋值DataFrame仅首次有效 我有一个通过随机过程返回整数的方法dotest(),遍历多索引DataFrametests的行时,用iterrows()返回的行对象给tests的tranche fails列赋值,但发现只有第一次循环赋值有效,后续循环里的赋值完全没生效。
现象展示
运行代码后输出如下:
tranche 0: 30000 tests tranche fails fails fails (%) uncertainty thickness sample size pval 3 30 0.05 6214 6214 20.71 0.47 5 30 0.05 3249 3249 10.83 0.36 tranche 1: 60000 tests tranche fails fails fails (%) uncertainty thickness sample size pval 3 30 0.05 6214 12428 20.71 0.33 5 30 0.05 3249 6498 10.83 0.25
后续所有循环中,tranche fails的数值都保持第一次的结果不变。但直接用tests.at[j, "tranche fails"]赋值就完全正常,我想搞懂为什么会出现“首次有效、后续失效”的情况。
精简代码
for i in range(0, tranches): for j, row in tests.iterrows(): row.at["tranche fails"] = dotest(tranche_size, j[1], j[0], j[2]) # tests.at[j, "tranche fails"] = dotest(tranche_size, j[1], j[0], j[2]) do_other_things() print(tests)
完整代码(含验证)
我用numpy.random.randint()替代dotest()测试,问题依然存在,说明不是dotest()的问题:
for i in range(0, tranches): testn = (i + 1) * tranche_size print(f"tranche {i}: {testn} tests") for j, row in tests.iterrows(): # row["tranche fails"] = np.random.randint(10) row.at["tranche fails"] = dotest(tranche_size, j[1], j[0], j[2]) # tests.at[j, "tranche fails"] = dotest(tranche_size, j[1], j[0], j[2]) tests["fails"] += tests["tranche fails"] p = tests["fails"]/testn tests["fails (%)"] = (100*p).round(2) # 用二项分布的正态近似计算不确定性 tests["uncertainty"] = (2*100*np.sqrt(p*(1-p)/testn)).round(2) try: prevtests.loc[tests.index, "fails"] += tests["tranche fails"] prevtests.loc[tests.index, "tests"] += tranche_size print(tests) prevtests.to_csv("Previous Tests.csv") except: pass
原因分析
这是因为iterrows()返回的row对象是原DataFrame的临时副本,不是对原数据的直接引用:
- 第一次循环时,
row副本的修改会被隐式同步回原DataFrame(这是pandas早期版本的设计模糊点,属于意外的弱关联行为); - 但当你第一次修改
tranche fails列后,该列的数据类型或内部存储结构发生了变化,后续iterrows()返回的row副本变成了完全独立的对象,修改副本不会再同步到原DataFrame。
简单说:第一次赋值时副本和原数据还存在弱关联,后续循环中这种关联被打破,修改副本自然不会影响原DataFrame。
解决方法
直接使用原DataFrame的索引赋值,避开iterrows()返回的副本问题:
- 用
tests.at[j, "tranche fails"](你代码中注释掉的那行),这是直接操作原DataFrame的高效方式; - 或者用
tests.loc[j, "tranche fails"],效果相同。
修改后的核心代码:
for i in range(0, tranches): for j, _ in tests.iterrows(): tests.at[j, "tranche fails"] = dotest(tranche_size, j[1], j[0], j[2]) do_other_things() print(tests)
内容的提问来源于stack exchange,提问作者Zoe Allen
相关产品推荐
相关产品推荐

