如何在Pandas中忽略空DataFrame的处理流程?
空DataFrame引发的AttributeError解决方法
原始DataFrame
No ITEM Location Class Data 14 Tiger__ Area 01 ***************************************PPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPP**********************************************************************************PPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPP 14 Tiger__ Area 02 **************************PPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPP*********************************************************************************PPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPP 14 Rabbit__ Area 01 **************************PPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPP*********************************************************************************PPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPP 14 Rabbit__ Area 02 **************************PPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPP*********************************************************************************PPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPPP
报错代码
Tigeritem = df[df['ITEM'].str.contains('Tiger_') == True] snakeitem = df[df['ITEM'].str.contains('Snake_') == True] Rabbititem = df[df['ITEM'].str.contains('Rabbit_') == True] realtiger = Tigeritem['Data'].astype(str).str.extractall('(\*+)')[0].str.len().loc[lambda x: x > 40].groupby(level=0).agg(list).str[1] realsnake = snakeitem['Data'].astype(str).str.extractall('(\*+)')[0].str.len().loc[lambda x: x > 40].groupby(level=0).agg(list).str[1] realrabbit = Rabbititem['Data'].astype(str).str.extractall('(\*+)')[0].str.len().loc[lambda x: x > 40].groupby(level=0).agg(list).str[1]
错误原因
由于DataFrame中没有包含Snake_的行,snakeitem是空DataFrame,执行realsnake相关代码时触发错误:
AttributeError: Can only use .str accessor with string values!
需求
跳过空DataFrame的处理流程,仅执行realtiger和realrabbit的处理,得到预期输出:
0 82 1 81 Name: 0, dtype: int64 2 81 3 81 Name: 0, dtype: int64
解决方案
封装一个通用处理函数,先判断DataFrame是否为空,非空时再执行提取逻辑,避免空DataFrame触发报错:
def process_item_data(item_df): if item_df.empty: return None return item_df['Data'].astype(str).str.extractall('(\*+)')[0].str.len().loc[lambda x: x > 40].groupby(level=0).agg(list).str[1] # 筛选各item对应的DataFrame Tigeritem = df[df['ITEM'].str.contains('Tiger_')] snakeitem = df[df['ITEM'].str.contains('Snake_')] Rabbititem = df[df['ITEM'].str.contains('Rabbit_')] # 仅处理非空的DataFrame realtiger = process_item_data(Tigeritem) realrabbit = process_item_data(Rabbititem) # 输出结果 print(realtiger) print(realrabbit)
说明
- 函数
process_item_data会先检查输入的DataFrame是否为空,为空则返回None,跳过后续的字符串操作 - 对有数据的
Tigeritem和Rabbititem正常执行提取逻辑,得到预期的长度值
内容的提问来源于stack exchange,提问作者888Seeme
相关产品推荐
相关产品推荐

