使用neattext提取hashtag报错TypeError: expected string or bytes-like object如何解决
报错原因
- 核心触发原因是
df['Tweets']列中存在非字符串/字节类型的值,包括空值NaN、None、数值类型等。nfx.extract_hashtags函数仅接收字符串或字节类输入,碰到不符合类型要求的输入就会抛出该TypeError。 - 可以先运行
df['Tweets'].apply(type).value_counts()排查该列的所有值类型,确认异常类型的分布。
解决方法
- 方案1:强制将整列转为字符串类型,空值会被转为字符串
'nan',不影响标签提取逻辑
代码:df['Tweets'].astype(str).apply(nfx.extract_hashtags) - 方案2:仅处理有效字符串行,跳过非字符串/空值记录
代码:# 筛选值为字符串类型的行 is_str_mask = df['Tweets'].apply(lambda x: isinstance(x, str)) df.loc[is_str_mask, 'hashtags'] = df.loc[is_str_mask, 'Tweets'].apply(nfx.extract_hashtags) - 方案3:封装函数增加异常捕获,碰到非法输入直接返回空列表
代码:def safe_extract(text): try: return nfx.extract_hashtags(text) except TypeError: return [] df['hashtags'] = df['Tweets'].apply(safe_extract)
内容的提问来源于stack exchange,提问作者becca David
相关产品推荐
相关产品推荐

