如何用非循环方式判断Pandas中当前值与前三个值均不相等?
判断DataFrame当前值与前三个值均不相等的矢量化实现
现有一段用于判断DataFrame中col1列当前值与前3个值是否相等的矢量化代码:
import pandas as pd data={'col1':[False, False, False, False,False, False, True, True]} df=pd.DataFrame(data,columns=['col1']) print(df) df['match'] = df.col1.eq(df.col1.shift(3)) print(df)
对应的输出结果为:
col1 0 False 1 False 2 False 3 False 4 False 5 False 6 True 7 True col1 match 0 False False 1 False False 2 False False 3 False True 4 False True 5 False True 6 True False 7 True False
现在需要实现当前值与前三个值均不相等的判断,要求不使用For循环,预期输出如下:
col1 match 0 False False 1 False False 2 False False 3 False False 4 False False 5 False False 6 True True 7 True False
解决方案
可以通过shift方法分别获取前1、2、3位的列数据,使用ne(不等于)运算符判断当前值与这三个值的关系,最后通过逻辑与&组合条件,再手动将前3行因缺少前置数据的match值设为False:
import pandas as pd data={'col1':[False, False, False, False,False, False, True, True]} df=pd.DataFrame(data,columns=['col1']) # 判断当前值与前1、2、3位的值均不相等 df['match'] = df['col1'].ne(df['col1'].shift(1)) & df['col1'].ne(df['col1'].shift(2)) & df['col1'].ne(df['col1'].shift(3)) # 前3行无足够前置数据,直接设为False df.loc[:2, 'match'] = False print(df)
也可以用更简洁的方式,通过concat合并当前列与前3个偏移列,再按行判断当前值与其他三个值均不相等:
import pandas as pd data={'col1':[False, False, False, False,False, False, True, True]} df=pd.DataFrame(data,columns=['col1']) # 合并当前列和前3个偏移列,逐行判断当前值与其他三个均不等 df['match'] = pd.concat( [df['col1'], df['col1'].shift(1), df['col1'].shift(2), df['col1'].shift(3)], axis=1 ).apply(lambda row: row[0] != row[1] and row[0] != row[2] and row[0] != row[3], axis=1) # 前3行手动置为False df.loc[:2, 'match'] = False print(df)
两种方法运行后都会得到预期输出:
col1 match 0 False False 1 False False 2 False False 3 False False 4 False False 5 False False 6 True True 7 True False
内容的提问来源于stack exchange,提问作者flintstone
相关产品推荐
相关产品推荐

