You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何用iterrows()行对象赋值首次有效后续失效?

问题:使用iterrows()返回的行对象赋值DataFrame仅首次有效

我有一个通过随机过程返回整数的方法dotest(),遍历多索引DataFrametests的行时,用iterrows()返回的行对象给tests的tranche fails列赋值,但发现只有第一次循环赋值有效,后续循环里的赋值完全没生效。

现象展示

运行代码后输出如下:

tranche 0: 30000 tests
                            tranche fails  fails  fails (%)  uncertainty
thickness sample size pval                                              
3         30          0.05           6214   6214      20.71         0.47
5         30          0.05           3249   3249      10.83         0.36
tranche 1: 60000 tests
                            tranche fails  fails  fails (%)  uncertainty
thickness sample size pval                                              
3         30          0.05           6214  12428      20.71         0.33
5         30          0.05           3249   6498      10.83         0.25

后续所有循环中,tranche fails的数值都保持第一次的结果不变。但直接用tests.at[j, "tranche fails"]赋值就完全正常,我想搞懂为什么会出现“首次有效、后续失效”的情况。

精简代码

for i in range(0, tranches):
    for j, row in tests.iterrows():
        row.at["tranche fails"] = dotest(tranche_size, j[1], j[0], j[2])
        # tests.at[j, "tranche fails"] = dotest(tranche_size, j[1], j[0], j[2])
    do_other_things()
    print(tests)

完整代码(含验证)

我用numpy.random.randint()替代dotest()测试,问题依然存在,说明不是dotest()的问题:

for i in range(0, tranches):
    testn = (i + 1) * tranche_size
    print(f"tranche {i}: {testn} tests")
    for j, row in tests.iterrows():
        # row["tranche fails"] = np.random.randint(10)
        row.at["tranche fails"] = dotest(tranche_size, j[1], j[0], j[2])
        # tests.at[j, "tranche fails"] = dotest(tranche_size, j[1], j[0], j[2])
    tests["fails"] += tests["tranche fails"]
    p = tests["fails"]/testn
    tests["fails (%)"] = (100*p).round(2)
    # 用二项分布的正态近似计算不确定性
    tests["uncertainty"] = (2*100*np.sqrt(p*(1-p)/testn)).round(2)
    try:
        prevtests.loc[tests.index, "fails"] += tests["tranche fails"]
        prevtests.loc[tests.index, "tests"] += tranche_size
        print(tests)
        prevtests.to_csv("Previous Tests.csv")
    except:
        pass

原因分析

这是因为iterrows()返回的row对象是原DataFrame的临时副本,不是对原数据的直接引用:

  • 第一次循环时,row副本的修改会被隐式同步回原DataFrame(这是pandas早期版本的设计模糊点,属于意外的弱关联行为);
  • 但当你第一次修改tranche fails列后,该列的数据类型或内部存储结构发生了变化,后续iterrows()返回的row副本变成了完全独立的对象,修改副本不会再同步到原DataFrame。

简单说:第一次赋值时副本和原数据还存在弱关联,后续循环中这种关联被打破,修改副本自然不会影响原DataFrame。


解决方法

直接使用原DataFrame的索引赋值,避开iterrows()返回的副本问题:

  • 用tests.at[j, "tranche fails"](你代码中注释掉的那行),这是直接操作原DataFrame的高效方式;
  • 或者用tests.loc[j, "tranche fails"],效果相同。

修改后的核心代码:

for i in range(0, tranches):
    for j, _ in tests.iterrows():
        tests.at[j, "tranche fails"] = dotest(tranche_size, j[1], j[0], j[2])
    do_other_things()
    print(tests)

内容的提问来源于stack exchange,提问作者Zoe Allen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 18:22:38