You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas递归.loc赋值除使用.unique()外还有哪些更优实现方法?

递归函数很难实现向量化,因为t时刻的每个输入都依赖t-1时刻的前序输入。

[下方问题已更新,补充了更复杂的计算示例$x_t = a x_{t-1} + b$。]

.loc返回不同数据类型的问题
import pandas
df1 = pandas.DataFrame({'year':range(2020,2024),'a':range(3,7)})
# 设置初始值
t0 = min(df1.year)
df1.loc[df1.year==t0, "x"] = 0

当等式右侧为pandas.core.series.Series类型时,该赋值逻辑无法生效:

for t in range (min(df1.year)+1, max(df1.year)+1):
    df1.loc[df1.year==t, "x"] = df1.loc[df1.year==t-1,"x"] + df1.loc[df1.year==t-1,"a"]
print(df1)
#    year  a    x
# 0  2020  3  0.0
# 1  2021  4  NaN
# 2  2022  5  NaN
# 3  2023  6  NaN
print(type(df1.loc[df1.year==t-1,"x"] + df1.loc[df1.year==t-1,"a"]))
# <class 'pandas.core.series.Series'>

当等式右侧为numpy数组时,赋值可正常生效:

for t in range (min(df1.year)+1, max(df1.year)+1):
    df1.loc[df1.year==t, "x"] = (df1.loc[df1.year==t-1,"x"] + df1.loc[df1.year==t-1,"a"]).unique()
    #break
print(df1)
#    year  a     x
# 0  2020  3   0.0
# 1  2021  4   3.0
# 2  2022  5   7.0
# 3  2023  6  12.0
print(type((df1.loc[df1.year==t-1,"x"] + df1.loc[df1.year==t-1,"a"]).unique()))
# <class 'numpy.ndarray'>

当使用year作为索引调用.loc()选取数据时,赋值可直接生效:

df2 = df1.set_index("year").copy()
# 设置初始值
df2.loc[df2.index.min(), "x"] = 0
for t in range (df2.index.min()+1, df2.index.max()+1):
    df2.loc[t, "x"] = df2.loc[t-1, "x"] + df2.loc[t-1,"a"]
    #break
print(df2)
#       a     x
# year
# 2020  3   0.0
# 2021  4   3.0
# 2022  5   7.0
# 2023  6  12.0
print(type(df2.loc[t-1, "x"] + df2.loc[t-1,"a"]))
# <class 'numpy.float64'>

问题解答

  1. 二者返回类型不同的原因
    核心是pandas的.loc取值规则差异:
  • 用df1.loc[df1.year==t-1,"x"]这种布尔掩码方式筛选时,pandas无法提前预判筛选结果只有1行,为了保证返回格式统一,不管匹配到多少行都会返回Series类型。两个Series相加结果还是Series,赋值时左侧目标行的索引和右侧Series的索引无法对齐,所以最终填充为NaN。
  • 用df2.loc[t-1, "x"]这种精确索引标签取值时,只要该标签在索引中唯一存在,pandas会直接返回标量值(这里为numpy.float64类型),标量赋值不需要索引对齐,所以可以直接写入目标位置。
  1. 不使用set_index的更优递归赋值方式
    比.unique()更稳定、可读性更好的方案是直接提取标量值,用.iloc[0]明确取出Series内的唯一值:
for t in range (min(df1.year)+1, max(df1.year)+1):
    prev_x = df1.loc[df1.year==t-1,"x"].iloc[0]
    prev_a = df1.loc[df1.year==t-1,"a"].iloc[0]
    df1.loc[df1.year==t, "x"] = prev_x + prev_a

如果是线性递归逻辑,也可以用expanding窗口实现向量化计算,性能比循环更优。


带乘法和加法组件的示例

实际业务场景的问题更为复杂,包含乘法和加法组合的递归计算逻辑:

import pandas
df3 = pandas.DataFrame({'year':range(2020,2024),'a':range(3,7), 'b':range(8,12)})
df3 = df3.set_index("year").copy()
# 设置初始值
df3.loc[df3.index.min(), "x"] = 0
for t in range (df3.index.min()+1, df3.index.max()+1):
    df3.loc[t, "x"] = df3.loc[t-1, "x"] * df3.loc[t-1, "a"] + df3.loc[t-1, "b"]
    #break
print(df3)

内容的提问来源于stack exchange,提问作者Paul Rougieux

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 09:15:04