You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用neattext提取hashtag报错TypeError: expected string or bytes-like object如何解决

报错原因
  • 核心触发原因是df['Tweets']列中存在非字符串/字节类型的值,包括空值NaN、None、数值类型等。nfx.extract_hashtags函数仅接收字符串或字节类输入,碰到不符合类型要求的输入就会抛出该TypeError。
  • 可以先运行df['Tweets'].apply(type).value_counts()排查该列的所有值类型,确认异常类型的分布。
解决方法
  • 方案1:强制将整列转为字符串类型,空值会被转为字符串'nan',不影响标签提取逻辑
    代码:
    df['Tweets'].astype(str).apply(nfx.extract_hashtags)
    
  • 方案2:仅处理有效字符串行,跳过非字符串/空值记录
    代码:
    # 筛选值为字符串类型的行
    is_str_mask = df['Tweets'].apply(lambda x: isinstance(x, str))
    df.loc[is_str_mask, 'hashtags'] = df.loc[is_str_mask, 'Tweets'].apply(nfx.extract_hashtags)
    
  • 方案3:封装函数增加异常捕获,碰到非法输入直接返回空列表
    代码:
    def safe_extract(text):
        try:
            return nfx.extract_hashtags(text)
        except TypeError:
            return []
    df['hashtags'] = df['Tweets'].apply(safe_extract)
    

内容的提问来源于stack exchange,提问作者becca David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 14:30:02