如何以Pythonic方式对Pandas DataFrame指定索引元素执行数学运算?
对Pandas DataFrame中部分索引元素执行数学运算的Pythonic写法
想知道对Pandas DataFrame里的部分索引元素做数学运算,正确的Pythonic方式是什么?试了几种方法,感觉繁琐还容易搞混,我的代码和运行结果如下:
df = pd.DataFrame({'x': [1, 2, 3, 4, 5, 6, 7, 9, ]}) df['y'] = df['x'] / 2 print(df) def printEpr(x): print(F"""End Point Rate: {epr} dy/dx""") epr = (df['y'][7] - df['y'][0]) / (df['x'][7] - df['x'][0]) printEpr(epr) epr = ( (df.iloc[[-1], [1]].values - df.iloc[[0], [1]].values) / (df.iloc[[-1], [0]].values - df.iloc[[1], [0]].values)) printEpr(epr) #epr = (df['y'] - df.iloc[0, 'y']) / (df.iloc[-1, 'x'] - df.iloc[0, 'x']) #printEpr(epr) yind = df.columns.get_loc('y') xind = df.columns.get_loc('x') epr = (df.iloc[-1, yind] - df.iloc[0, yind]) / (df.iloc[-1, xind] - df.iloc[0, xind]) printEpr(epr) #epr = (df.iloc[-1, 'y'] - df.iloc[0, 'y']) / (df.iloc[-1, 'x'] - df.iloc[0, 'x']) #printEpr(epr)
运行结果:
x y 0 1 0.5 1 2 1.0 2 3 1.5 3 4 2.0 4 5 2.5 5 6 3.0 6 7 3.5 7 9 4.5 End Point Rate: 0.5 dy/dx End Point Rate: [[0.57142857]] dy/dx End Point Rate: 0.5 dy/dx
简洁且Pythonic的写法
这里推荐几种更清晰、简洁的实现方式,避免冗余代码:
- 先提取首尾行,再通过列名取值
这种写法可读性最强,逻辑清晰:
first = df.iloc[0] last = df.iloc[-1] epr = (last['y'] - first['y']) / (last['x'] - first['x']) printEpr(epr) # 输出 0.5 dy/dx
- 直接用列+位置索引取值
如果想一行写完,也可以这样写(注意:链式索引df['y'].iloc[-1]取值时没问题,赋值时可能触发SettingWithCopyWarning,但取值场景下安全):
epr = (df['y'].iloc[-1] - df['y'].iloc[0]) / (df['x'].iloc[-1] - df['x'].iloc[0])
- 用
loc按索引标签+列名取值
如果你的DataFrame用的是自定义标签索引,这种方式更通用:
# 这里索引是默认整数,所以df.index[-1]就是最后一行的标签7 epr = (df.loc[df.index[-1], 'y'] - df.loc[df.index[0], 'y']) / (df.loc[df.index[-1], 'x'] - df.loc[df.index[0], 'x'])
你代码里的问题说明
- 第二个计算结果错误是因为分母用了
df.iloc[[1], [0]](取第2行的x值2),而不是df.iloc[0, 0],导致分母变成9-2=7,结果出现偏差。 - 被注释的
df.iloc[-1, 'y']写法错误,iloc是基于位置的索引,不能直接传列名,要用列的位置索引或者换loc。
内容的提问来源于stack exchange,提问作者Dave X
相关产品推荐
相关产品推荐

