使用df.loc赋值失败:为何条件匹配的df.loc赋值不生效?
问题:df.loc赋值后目标列变为NaN的排查与解决
我使用df.loc筛选行进行赋值操作,确认两组筛选结果行数一致,但执行赋值后目标列new的值全部变成了NaN。相关代码及执行结果如下:
代码示例
print(df.loc[df['type'] == "OPEN", "A"]) print(df.loc[df['type'] == "CLOSE", "new"]) df.loc[df['type'] == "CLOSE", "new"] = df.loc[df['type'] == "OPEN", "A"] print(df.loc[df['type'] == "OPEN", "A"]) print(df.loc[df['type'] == "CLOSE", "new"])
执行结果
19 25790.00092 42 25565.00053 Name: trade_amount, dtype: float64 26 0 45 0 Name: new, dtype: int64 19 25790.00092 42 25565.00053 Name: trade_amount, dtype: float64 26 NaN 45 NaN
原因分析
核心问题是索引不匹配。从执行结果可以看到:
- 筛选
type=="OPEN"得到的行索引是19、42 - 筛选
type=="CLOSE"得到的行索引是26、45
Pandas的赋值操作默认会按索引对齐,当右侧Series的索引在左侧目标行中不存在时,对应位置就会被填充为NaN。
解决方法
方法1:提取纯值(忽略索引)
直接取出右侧Series的数值数组,跳过索引对齐逻辑:
df.loc[df['type'] == "CLOSE", "new"] = df.loc[df['type'] == "OPEN", "A"].values
方法2:重置索引对齐
将右侧Series的索引重置为连续整数,与左侧目标行的位置匹配:
open_values = df.loc[df['type'] == "OPEN", "A"].reset_index(drop=True) df.loc[df['type'] == "CLOSE", "new"] = open_values
方法3:使用np.where直接赋值
通过条件判断直接赋值,绕开索引对齐问题:
import numpy as np df['new'] = np.where(df['type'] == "CLOSE", df.loc[df['type'] == "OPEN", "A"].values, df['new'])
内容的提问来源于stack exchange,提问作者whatever
相关产品推荐
相关产品推荐

