You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:解决Jupyter中运行文本预处理代码出现的'float'无strip属性错误

解决方法

错误原因

报错是因为reviews.comment列里混有float类型的数据(大概率是缺失值NaN,pandas里字符串列出现缺失值时会自动转为float类型),而float对象没有strip()方法。

具体修复代码

方案一(推荐,用pandas原生方法更高效)

# 先把comment列统一转为字符串,同时把缺失值(NaN)替换为空字符串
reviews['comment'] = reviews['comment'].astype(str).fillna('')
# 执行去首尾空格+过滤空字符串的操作
cleaned_reviews = [comment.strip() for comment in reviews['comment']]
cleaned_reviews = [comment for comment in cleaned_reviews if comment]
# 查看前10条结果
cleaned_reviews[:10]

方案二(列表推导式里直接做类型判断)

如果不想修改原DataFrame,也可以在循环里先判断元素类型:

cleaned_reviews = []
for comment in reviews.comment:
    # 只处理字符串类型的评论,非字符串(比如NaN)直接跳过
    if isinstance(comment, str):
        stripped_comment = comment.strip()
        # 过滤掉去空格后为空的内容
        if stripped_comment:
            cleaned_reviews.append(stripped_comment)
# 查看前10条结果
cleaned_reviews[:10]

为啥之前的尝试没用?

  • 单纯移除strip():还是会保留float类型的元素,后续操作可能继续报错;
  • 盲目移除float:如果没正确判断元素类型就过滤,要么没过滤干净,要么误删了有效数据。

内容的提问来源于stack exchange,提问作者user21064497

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 02:45:35