You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas处理Excel做NLP文本清洗报float对象不可迭代错误求解

报错根因

你遇到的TypeError: 'float' object is not iterable由两个问题共同导致:

  1. 读取的Excel文件中text_4文本列存在空白单元格,pandas会将这类空值默认解析为float类型的NaN。clean_text函数中遍历字符的逻辑默认输入是可迭代的字符串,传入浮点型NaN时无法执行迭代,直接抛出错误。部分文件可正常运行,是因为这些文件的文本列恰好没有缺失值。
  2. 原代码存在缩进错误:data['text_5']的赋值语句被错误缩进在clean_text函数内部,运行时会触发语法错误。
修复步骤
  • 在clean_text函数入口增加类型校验兜底,将所有非字符串类型的输入统一转换为空字符串,从根源避免遍历非可迭代对象的问题
  • 调整错误缩进的代码,将列处理逻辑移到函数外部
  • 针对单文件50万行的大数据量,可在读取数据后先统计文本列空值数量,避免跑数中途崩溃浪费时间
修正后完整代码
import nltk
import pandas as pd
import re
import string
from nltk.corpus import stopwords

# 提前下载依赖,避免首次运行报错
nltk.download('wordnet')
nltk.download('stopwords')
wn = nltk.WordNetLemmatizer()
ps = nltk.PorterStemmer()

pd.set_option("display.max_colwidth", 200)
stop_word = set(stopwords.words("english"))
excel_path = r"C:xxxx-BTC 21-30-4.xlsx"
data = pd.read_excel(excel_path)
data.columns = ["#", "id", "date", "Name", "text_4"]

# 可选:提前排查空值
print(f"文本列空值总数:{data['text_4'].isna().sum()}")
# 可选:直接丢弃空文本行减少无效计算
# data = data.dropna(subset=['text_4']).reset_index(drop=True)

def clean_text(txt):
    # 类型兜底:非字符串输入统一转为空字符串
    if not isinstance(txt, str):
        txt = ""
    txt = "".join([c for c in txt if c not in string.punctuation])
    tokens = re.split('\W+', txt)
    # 过滤空字符
    txt = [word for word in tokens if word not in stop_word and word.strip() != ""]
    return txt

data['text_5'] = data['text_4'].apply(lambda x: clean_text(x))

def lemmatization(token_txt):
    text = [wn.lemmatize(word) for word in token_txt]
    return text

data["text_6"] = data['text_5'].apply(lambda x: lemmatization(x))

print(data.head())
大文件处理优化建议
  • 读取Excel时通过usecols参数指定只加载需要的5列,不要加载全量无关列,可大幅降低内存占用
  • 如果内存吃紧,可将数据分块读取处理,避免单批次数据量过大触发内存溢出
  • 词形还原步骤如果速度过慢,可考虑替换为多进程版本的apply,50万行数据处理速度可提升3-5倍

内容的提问来源于stack exchange,提问作者Andrew Grice

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 13:24:22