You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Spacy文本预处理报错:预期输入为字符串/Doc/bytes,却获<float>类型

报错原因与解决方法

报错根源

你的数据集text列中存在NaN值(在pandas里NaN以float类型存储),而spaCy的nlp()方法仅接受字符串、Doc或字节类型输入,因此触发ValueError [E1041]。

解决步骤

1. 清理数据中的NaN值

先检查text列的NaN数量:

print(df_balanced['text'].isna().sum())

根据需求选择处理方式:

  • 删除含NaN的行(适合NaN占比低的情况):
    df_balanced = df_balanced.dropna(subset=['text'])
    
  • 用空字符串填充NaN(保留所有行):
    df_balanced['text'] = df_balanced['text'].fillna('')
    

2. 给预处理函数添加容错机制

修改tf_preprocess函数,提前判断输入类型,避免单个异常中断整个apply过程:

# Further Preprocessing
# remove stopwords and puntuations
tf_nlp = spacy.load("en_core_web_sm")
def tf_preprocess(text):
    # 非字符串类型直接返回空字符串
    if not isinstance(text, str):
        return ""
    doc_15 = tf_nlp(text)
    filtered_data = []
    for token in doc_15:
        if token.is_stop or token.is_punct:
            continue
        filtered_data.append(token.lemma_)
    return " ".join(filtered_data)

建议优先做数据清理,再运行预处理,这样能从根源避免问题。

内容的提问来源于stack exchange,提问作者Rumple Rak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 16:13:24