Pandas递归.loc赋值除使用.unique()外还有哪些更优实现方法?
递归函数很难实现向量化,因为t时刻的每个输入都依赖t-1时刻的前序输入。
[下方问题已更新,补充了更复杂的计算示例$x_t = a x_{t-1} + b$。]
.loc返回不同数据类型的问题
import pandas df1 = pandas.DataFrame({'year':range(2020,2024),'a':range(3,7)}) # 设置初始值 t0 = min(df1.year) df1.loc[df1.year==t0, "x"] = 0
当等式右侧为pandas.core.series.Series类型时,该赋值逻辑无法生效:
for t in range (min(df1.year)+1, max(df1.year)+1): df1.loc[df1.year==t, "x"] = df1.loc[df1.year==t-1,"x"] + df1.loc[df1.year==t-1,"a"] print(df1) # year a x # 0 2020 3 0.0 # 1 2021 4 NaN # 2 2022 5 NaN # 3 2023 6 NaN print(type(df1.loc[df1.year==t-1,"x"] + df1.loc[df1.year==t-1,"a"])) # <class 'pandas.core.series.Series'>
当等式右侧为numpy数组时,赋值可正常生效:
for t in range (min(df1.year)+1, max(df1.year)+1): df1.loc[df1.year==t, "x"] = (df1.loc[df1.year==t-1,"x"] + df1.loc[df1.year==t-1,"a"]).unique() #break print(df1) # year a x # 0 2020 3 0.0 # 1 2021 4 3.0 # 2 2022 5 7.0 # 3 2023 6 12.0 print(type((df1.loc[df1.year==t-1,"x"] + df1.loc[df1.year==t-1,"a"]).unique())) # <class 'numpy.ndarray'>
当使用year作为索引调用.loc()选取数据时,赋值可直接生效:
df2 = df1.set_index("year").copy() # 设置初始值 df2.loc[df2.index.min(), "x"] = 0 for t in range (df2.index.min()+1, df2.index.max()+1): df2.loc[t, "x"] = df2.loc[t-1, "x"] + df2.loc[t-1,"a"] #break print(df2) # a x # year # 2020 3 0.0 # 2021 4 3.0 # 2022 5 7.0 # 2023 6 12.0 print(type(df2.loc[t-1, "x"] + df2.loc[t-1,"a"])) # <class 'numpy.float64'>
问题解答
- 二者返回类型不同的原因
核心是pandas的.loc取值规则差异:
- 用
df1.loc[df1.year==t-1,"x"]这种布尔掩码方式筛选时,pandas无法提前预判筛选结果只有1行,为了保证返回格式统一,不管匹配到多少行都会返回Series类型。两个Series相加结果还是Series,赋值时左侧目标行的索引和右侧Series的索引无法对齐,所以最终填充为NaN。 - 用
df2.loc[t-1, "x"]这种精确索引标签取值时,只要该标签在索引中唯一存在,pandas会直接返回标量值(这里为numpy.float64类型),标量赋值不需要索引对齐,所以可以直接写入目标位置。
- 不使用set_index的更优递归赋值方式
比.unique()更稳定、可读性更好的方案是直接提取标量值,用.iloc[0]明确取出Series内的唯一值:
for t in range (min(df1.year)+1, max(df1.year)+1): prev_x = df1.loc[df1.year==t-1,"x"].iloc[0] prev_a = df1.loc[df1.year==t-1,"a"].iloc[0] df1.loc[df1.year==t, "x"] = prev_x + prev_a
如果是线性递归逻辑,也可以用expanding窗口实现向量化计算,性能比循环更优。
带乘法和加法组件的示例
实际业务场景的问题更为复杂,包含乘法和加法组合的递归计算逻辑:
import pandas df3 = pandas.DataFrame({'year':range(2020,2024),'a':range(3,7), 'b':range(8,12)}) df3 = df3.set_index("year").copy() # 设置初始值 df3.loc[df3.index.min(), "x"] = 0 for t in range (df3.index.min()+1, df3.index.max()+1): df3.loc[t, "x"] = df3.loc[t-1, "x"] * df3.loc[t-1, "a"] + df3.loc[t-1, "b"] #break print(df3)
内容的提问来源于stack exchange,提问作者Paul Rougieux
相关产品推荐
相关产品推荐

