pandas处理Excel做NLP文本清洗报float对象不可迭代错误求解
报错根因
你遇到的TypeError: 'float' object is not iterable由两个问题共同导致:
- 读取的Excel文件中
text_4文本列存在空白单元格,pandas会将这类空值默认解析为float类型的NaN。clean_text函数中遍历字符的逻辑默认输入是可迭代的字符串,传入浮点型NaN时无法执行迭代,直接抛出错误。部分文件可正常运行,是因为这些文件的文本列恰好没有缺失值。 - 原代码存在缩进错误:
data['text_5']的赋值语句被错误缩进在clean_text函数内部,运行时会触发语法错误。
修复步骤
- 在
clean_text函数入口增加类型校验兜底,将所有非字符串类型的输入统一转换为空字符串,从根源避免遍历非可迭代对象的问题 - 调整错误缩进的代码,将列处理逻辑移到函数外部
- 针对单文件50万行的大数据量,可在读取数据后先统计文本列空值数量,避免跑数中途崩溃浪费时间
修正后完整代码
import nltk import pandas as pd import re import string from nltk.corpus import stopwords # 提前下载依赖,避免首次运行报错 nltk.download('wordnet') nltk.download('stopwords') wn = nltk.WordNetLemmatizer() ps = nltk.PorterStemmer() pd.set_option("display.max_colwidth", 200) stop_word = set(stopwords.words("english")) excel_path = r"C:xxxx-BTC 21-30-4.xlsx" data = pd.read_excel(excel_path) data.columns = ["#", "id", "date", "Name", "text_4"] # 可选:提前排查空值 print(f"文本列空值总数:{data['text_4'].isna().sum()}") # 可选:直接丢弃空文本行减少无效计算 # data = data.dropna(subset=['text_4']).reset_index(drop=True) def clean_text(txt): # 类型兜底:非字符串输入统一转为空字符串 if not isinstance(txt, str): txt = "" txt = "".join([c for c in txt if c not in string.punctuation]) tokens = re.split('\W+', txt) # 过滤空字符 txt = [word for word in tokens if word not in stop_word and word.strip() != ""] return txt data['text_5'] = data['text_4'].apply(lambda x: clean_text(x)) def lemmatization(token_txt): text = [wn.lemmatize(word) for word in token_txt] return text data["text_6"] = data['text_5'].apply(lambda x: lemmatization(x)) print(data.head())
大文件处理优化建议
- 读取Excel时通过
usecols参数指定只加载需要的5列,不要加载全量无关列,可大幅降低内存占用 - 如果内存吃紧,可将数据分块读取处理,避免单批次数据量过大触发内存溢出
- 词形还原步骤如果速度过慢,可考虑替换为多进程版本的apply,50万行数据处理速度可提升3-5倍
内容的提问来源于stack exchange,提问作者Andrew Grice
相关产品推荐
相关产品推荐

