求助:解决Jupyter中运行文本预处理代码出现的'float'无strip属性错误
解决方法
错误原因
报错是因为reviews.comment列里混有float类型的数据(大概率是缺失值NaN,pandas里字符串列出现缺失值时会自动转为float类型),而float对象没有strip()方法。
具体修复代码
方案一(推荐,用pandas原生方法更高效)
# 先把comment列统一转为字符串,同时把缺失值(NaN)替换为空字符串 reviews['comment'] = reviews['comment'].astype(str).fillna('') # 执行去首尾空格+过滤空字符串的操作 cleaned_reviews = [comment.strip() for comment in reviews['comment']] cleaned_reviews = [comment for comment in cleaned_reviews if comment] # 查看前10条结果 cleaned_reviews[:10]
方案二(列表推导式里直接做类型判断)
如果不想修改原DataFrame,也可以在循环里先判断元素类型:
cleaned_reviews = [] for comment in reviews.comment: # 只处理字符串类型的评论,非字符串(比如NaN)直接跳过 if isinstance(comment, str): stripped_comment = comment.strip() # 过滤掉去空格后为空的内容 if stripped_comment: cleaned_reviews.append(stripped_comment) # 查看前10条结果 cleaned_reviews[:10]
为啥之前的尝试没用?
- 单纯移除
strip():还是会保留float类型的元素,后续操作可能继续报错; - 盲目移除float:如果没正确判断元素类型就过滤,要么没过滤干净,要么误删了有效数据。
内容的提问来源于stack exchange,提问作者user21064497
相关产品推荐
相关产品推荐

