修复Python筛选学历代码AttributeError: 'float' object has no attribute 'find'报错
错误原因
你的数据集Last_degree列中存在空值(NaN),Pandas中的NaN属于float类型,不存在字符串类的find()方法,因此遍历到空值时就会触发当前报错。
修复代码
调整degree函数,在执行字符串匹配前先做空值和类型校验,同时修正你原代码末尾的变量名拼写错误:
import pandas as pd edu = Edu_data # 调整后的学历识别函数 def degree(x): # 空值或非字符串类型直接返回0 if pd.isna(x) or not isinstance(x, str): return 0 if x.find('Bachelor') != -1 or x.find("Bachelor's") != -1 or x.find('BS') != -1 or x.find('bs') != -1: return 1 if x.find('Master') != -1 or x.find("Master's") != -1 or x.find('M.S') != -1 or x.find('MS') != -1 or x.find('MPhil') != -1 or x.find('MBA') != -1 or x.find('MicroMasters') != -1 or x.find('MSc') != -1 or x.find('MSCS') !=-1 or x.find('MSDS')!=-1: return 2 if x.find('PhD') != -1 or x.find('P.hd') != -1 or x.find('Ph.D') != -1 or x.find('ph.d') != -1: return 3 else: return 0 # 生成学历分类列 edu['deg'] = edu['Last_degree'].apply(degree) # 原代码此处变量名拼写错误,已修正 print(edu)
可选性能优化方案
你可以直接用Pandas矢量化字符串操作实现相同逻辑,无需自定义函数遍历,运行效率更高:
edu['deg'] = 0 # 匹配本科学历赋值1,na参数自动处理空值 edu.loc[edu['Last_degree'].str.contains(r'Bachelor|BS|bs|Bs|Bachelors', na=False), 'deg'] = 1 # 匹配硕士学历赋值2 edu.loc[edu['Last_degree'].str.contains(r"Master|MS|M\.S|MPhil|MBA|MicroMasters|MSc|MSCS|MSDS", na=False), 'deg'] = 2 # 匹配博士学历赋值3 edu.loc[edu['Last_degree'].str.contains(r'PhD|Ph\.D|P\.hd|ph\.d', na=False), 'deg'] = 3
内容的提问来源于stack exchange,提问作者Basharat Hussaain
相关产品推荐
相关产品推荐

