如何在Pandas的lambda函数中查找当前行前最近的long/short记录?
问题描述
处理Pandas数据时,需要在每行的lambda处理逻辑中,获取当前行之前最近的openingMethod为long或short的记录,并基于该行信息进行后续判断。例如处理行4(openingMethod为hold)时,需找到行1的short记录。
示例DataFrame(dfFinal):
openingMethod spread 0 long 10 1 short -10 2 hold 110 3 hold -20 4 hold -100 5 long 150 6 short -250 7 hold 210 8 hold -120 9 hold 130
现有代码中,_computeIfMakeMoney方法无法实现获取最近long/short记录的逻辑,需要解决该问题。
解决方案
推荐提前对DataFrame进行预处理,生成一个存储每行对应最近long/short值的列,这种方式利用Pandas矢量化操作,效率远高于逐行查找。
步骤说明
- 生成最近值列:创建临时列,仅保留
openingMethod为long或short的行值,其余设为NaN;再通过ffill()向前填充,将最近的有效值填充到后续的hold行中。 - 在处理逻辑中调用:在
_computeIfMakeMoney方法中,直接通过当前行索引读取预处理好的最近值列即可。
修改后的完整代码
import pandas as pd import numpy as np class DataProcessor: def __init__(self): self.dfFinal = self.setTempData() # 预处理:生成最近的long/short记录列 self._preprocess_nearest_opening() self.tempProfitList = [] self.tempAllProfit: float = 0.0 @staticmethod def setTempData() -> pd.DataFrame: df = pd.DataFrame(columns=['openingMethod', 'spread']) df.loc[len(df)] = {"openingMethod": "long", "spread": 10} df.loc[len(df)] = {"openingMethod": "short", "spread": -10} df.loc[len(df)] = {"openingMethod": "hold", "spread": 110} df.loc[len(df)] = {"openingMethod": "hold", "spread": -20} df.loc[len(df)] = {"openingMethod": "hold", "spread": -100} df.loc[len(df)] = {"openingMethod": "long", "spread": 150} df.loc[len(df)] = {"openingMethod": "short", "spread": -250} df.loc[len(df)] = {"openingMethod": "hold", "spread": 210} df.loc[len(df)] = {"openingMethod": "hold", "spread": -120} df.loc[len(df)] = {"openingMethod": "hold", "spread": 130} return df def _preprocess_nearest_opening(self): # 创建临时列,仅保留long/short值,其余为NaN temp_col = self.dfFinal['openingMethod'].where( self.dfFinal['openingMethod'].isin(['long', 'short']) ) # 向前填充,得到每行最近的long/short值 self.dfFinal['preNearestOpeningMethod'] = temp_col.ffill() def setIfMakeMoney(self): self.dfFinal["ifMakeMoney"] = self.dfFinal.apply( lambda x: self._computeIfMakeMoney(x.name, x['spread']), axis=1) def _computeIfMakeMoney(self, openingMethodIndex, spread): # 直接读取预处理好的最近long/short值 preNearestOpeningMethod = self.dfFinal.loc[openingMethodIndex, 'preNearestOpeningMethod'] if not np.isnan(spread): self.tempProfitList.append(spread) # 简化求和逻辑,用sum替代循环 self.tempAllProfit = sum(s for s in self.tempProfitList if not np.isnan(s)) if self.tempAllProfit > 0 and preNearestOpeningMethod == "long": self.tempProfitList = [] return f"Get Money is {self.tempAllProfit}" elif self.tempAllProfit < 0 and preNearestOpeningMethod == "short": self.tempProfitList = [] return f"Get Money is {-self.tempAllProfit}" elif self.tempAllProfit < 0 and preNearestOpeningMethod == "long": self.tempProfitList = [] return f"Loss Money is {self.tempAllProfit}" elif self.tempAllProfit > 0 and preNearestOpeningMethod == "short": self.tempProfitList = [] return f"Loss Money is {-self.tempAllProfit}" else: return "holding" def main(): DP = DataProcessor() DP.setIfMakeMoney() print(DP.dfFinal) if __name__ == "__main__": main()
额外优化说明
- 原代码中循环求和的逻辑被替换为
sum(s for s in self.tempProfitList if not np.isnan(s)),简化代码同时提升效率。 - 字符串拼接改用f-string,更简洁易读。
内容的提问来源于stack exchange,提问作者陈哲珂
相关产品推荐
相关产品推荐

