DataFrame列字符串匹配报错:'str'对象无'str'属性排查问询
问题与解决
问题场景
我在处理DataFrame时写了个函数,想检查Campaign列每行是否包含“Bathroom”相关字符串(比如'Bathrooms'、'Bathrooms - Des Moines'这类),如果包含就在新列Product里填“Bathroom”。但运行代码时触发了属性错误:'str' object has no attribute 'str',代码如下:
def product(x): bathroom = x['Campaign'].str.contains('Bathroom') if bathroom == True: return 'Bathroom' df['Product'] = df.apply(product, axis = 1)
错误原因
当你用df.apply(..., axis=1)的时候,传给product函数的x是DataFrame的单行数据(Series对象),而x['Campaign']取到的是这一行里Campaign列的单个字符串值,不是Series。.str是pandas给Series/列批量处理字符串的专属属性,单个Python字符串本身没有这个属性,所以直接报错。
解决方案
方案1:改用Python字符串原生判断
把.str.contains换成Python原生的字符串包含判断即可:
def product(x): if 'Bathroom' in x['Campaign']: return 'Bathroom' df['Product'] = df.apply(product, axis=1)
方案2:用pandas向量化操作(更高效)
如果数据量较大,apply的循环效率很低,直接对整个Campaign列用.str.contains做批量判断,再赋值:
# 方法一:用map映射 df['Product'] = df['Campaign'].str.contains('Bathroom').map({True: 'Bathroom', False: None}) # 方法二:用numpy.where更直观 import numpy as np df['Product'] = np.where(df['Campaign'].str.contains('Bathroom'), 'Bathroom', None)
补充提示
- 方案1适合需要在函数里添加更多复杂分支逻辑的场景;
- 方案2是pandas原生的向量化操作,比
apply快数倍到数十倍,数据量大时优先选这个。
内容的提问来源于stack exchange,提问作者internet joe
相关产品推荐
相关产品推荐

