如何在DataFrame中基于当前行与前一行数据创建新列?
问题:基于DataFrame当前行与前一行数据生成新列
我需要从包含纬度(lats)、经度(lons)、海拔(els)、时间(ts,单位秒)的元组生成整理后的文件,已经将元组转为DataFrame。现在要添加名为stepsizes的新列,该列需要调用一个依赖当前行与前一行经纬度的函数:
def stepsize(lat1, long1, lat2, long2): # 两点距离计算逻辑
尝试用df.apply()结合lambda函数实现时,因局部变量无法在行之间传递,触发"引用未赋值"报错:
def newdist_row(row): if row["time"] < 72889: rowprevlat = row['latitude'] rowprevlong = row['longitude'] rowprevtime = row['time'] return np.nan else: dist = stepsize_feet(row['latitude'], row['longitude'], rowprevlat, rowprevlong) rowprevlat = row['latitude'] rowprevlong = row['longitude'] rowprevtime = row['time'] return dist contents.apply(lambda row: newdist_row(row), axis=1)
也试过循环遍历行,但出现"试图在DataFrame切片副本上设置值"的警告,均未成功。
解决方法:使用pandas的shift()方法
shift()是处理前后行数据的标准方案,无需手动维护变量,效率远高于apply()或循环:
- 生成前一行的经纬度临时列:
# 获取前一行的纬度、经度数据 contents['prev_latitude'] = contents['latitude'].shift(1) contents['prev_longitude'] = contents['longitude'].shift(1)
- 用
loc定位赋值,避免切片副本警告并计算目标列:
import numpy as np # 定义计算条件:时间≥72889时计算距离,否则为NaN mask = contents['time'] >= 72889 # 对符合条件的行调用距离计算函数 contents.loc[mask, 'stepsizes'] = contents.loc[mask].apply( lambda x: stepsize_feet(x['latitude'], x['longitude'], x['prev_latitude'], x['prev_longitude']), axis=1 ) # 不符合条件的行设置为NaN contents.loc[~mask, 'stepsizes'] = np.nan
- (可选)优化:将距离函数改为向量化实现,彻底抛弃
apply提升效率:
def stepsize_feet_vec(lats1, lons1, lats2, lons2): # 改为支持Series输入的向量化距离计算逻辑 pass contents['stepsizes'] = np.where( contents['time'] >=72889, stepsize_feet_vec(contents['latitude'], contents['longitude'], contents['prev_latitude'], contents['prev_longitude']), np.nan )
- 清理临时列(可选):
contents.drop(['prev_latitude', 'prev_longitude'], axis=1, inplace=True)
内容的提问来源于stack exchange,提问作者Mike
相关产品推荐
相关产品推荐

