如何在Pandas DataFrame中检测列表内U后是否跟非U/None值?
问题详情
创建的DataFrame代码
import pandas as pd import numpy as np ds = { 'col1' : [ ['U', 'U', 'U', 'U', 'U', 1, 0, 0, 0, 'U','U', None], [6, 5, 4, 3, 2], [0, 0, 0, 'U', 'U'], [0, 1, 'U', 'U', 'U'], [0, 'U', 'U', 'U', None] ] } df = pd.DataFrame(data=ds)
DataFrame内容
col1 0 [U, U, U, U, U, 1, 0, 0, 0, U, U, None] 1 [6, 5, 4, 3, 2] 2 [0, 0, 0, U, U] 3 [0, 1, U, U, U] 4 [0, U, U, U, None]
需求
对col1的每一行,检查列表中是否存在元素'U'的下一个元素既不是'U'也不是None,若是则新建iCount列赋值为1,否则为0。预期结果:
col1 iCount 0 [U, U, U, U, U, 1, 0, 0, 0, U, U, None] 1 1 [6, 5, 4, 3, 2] 0 2 [0, 0, 0, U, U] 0 3 [0, 1, U, U, U] 0 4 [0, U, U, U, None] 0
尝试的代码及错误结果
尝试的代码:
col5 = np.array(df['col1']) for i in range(len(df)): iCount = 0 for j in range(len(col5[i])-1): print(col5[i][j]) if((col5[i][j] == "U") & ((col5[i][j+1] != None) & (col5[i][j+1] != "U"))): iCount += 1 else: iCount = iCount
错误结果:
col1 iCount 0 [U, U, U, U, U, 1, 0, 0, 0, U, U, None] 0 1 [6, 5, 4, 3, 2] 0 2 [0, 0, 0, U, U] 0 3 [0, 1, U, U, U] 0 4 [0, U, U, U, None] 0
错误原因分析
- None比较方式错误:Python中判断是否为None必须用
is not None,而非!= None。!=会调用对象的__ne__方法,可能出现不符合预期的结果;is是判断对象身份,None是单例,用is才是正确逻辑。 - 未将结果赋值回DataFrame:循环计算完
iCount后,没有把值存入DataFrame的iCount列,导致所有行的iCount默认是0。 - 逻辑冗余:需求是判断“是否存在”符合条件的情况,只要找到一次就可以将
iCount设为1,无需累加,还能提前跳出循环提升效率。
解决方案
方案1:修正原循环代码
# 先初始化iCount列 df['iCount'] = 0 col5 = df['col1'].tolist() for i in range(len(df)): current_list = col5[i] count = 0 # 遍历到倒数第二个元素 for j in range(len(current_list)-1): if current_list[j] == 'U': next_val = current_list[j+1] # 正确判断None和非U if next_val is not None and next_val != 'U': count = 1 break # 找到一次就跳出循环,无需继续检查 df.loc[i, 'iCount'] = count
方案2:使用Pandas apply函数(更简洁)
利用apply对每行的列表进行处理,用zip将列表和其偏移一位的列表配对,检查是否存在符合条件的配对:
def check_condition(lst): # 遍历相邻元素对 for curr, next_val in zip(lst, lst[1:]): if curr == 'U' and next_val is not None and next_val != 'U': return 1 return 0 df['iCount'] = df['col1'].apply(check_condition)
内容的提问来源于stack exchange,提问作者Giampaolo Levorato
相关产品推荐
相关产品推荐

