Python处理Twitter数据时,对列应用清洗函数报错:object of type 'float' has no len()
Python处理Twitter数据时,对列应用清洗函数报错:object of type 'float' has no len()
嘿,我看你遇到的这个问题挺常见的,咱们来一步步拆解解决哈!
问题原因分析
你单独处理某一条评论时没问题,但给整列用apply就报错object of type 'float' has no len(),这十有八九是你的comment列里藏着空值(NaN)——在pandas里,NaN的类型是float,当函数遍历到这些空值时,后续的字符串处理步骤(比如BeautifulSoup解析、正则替换)就会因为传入的不是字符串而触发错误。
解决方案
1. 先处理数据里的空值
你可以根据需求选下面两种方式:
- 方式一:直接删除含空值的行(如果空值占比低,不影响后续分析的话)
# 只删除comment列有空值的行 twitter_data = twitter_data.dropna(subset=['comment'])
- 方式二:把空值替换成空字符串(想保留所有行,后续清洗后对应结果为空)
twitter_data['comment'] = twitter_data['comment'].fillna('')
2. 给清洗函数加防御性检查+效率优化
即使处理了空值,也可以在函数里加一层保险,同时把重复初始化的对象移到函数外提升效率:
import pandas as pd from bs4 import BeautifulSoup import re import nltk from nltk.corpus import stopwords from nltk.stem import SnowballStemmer, WordNetLemmatizer # 只初始化一次,避免每次调用函数重复执行 SW = set(stopwords.words('english')) # 转成集合,查找速度更快 SS_stem = SnowballStemmer(language='english') word_lemmitize = WordNetLemmatizer() def data_clean_pipeline(text): # 先处理非字符串的情况,比如NaN或者float类型 if not isinstance(text, str): # NaN转空字符串,其他非字符串类型转成字符串 text = '' if pd.isna(text) else str(text) # 原有清洗逻辑 text = str(BeautifulSoup(text).get_text()) text = re.sub("[^a-zA-Z]", " ", text) text = text.lower() text = nltk.word_tokenize(text) text = [t for t in text if t not in SW] text = [SS_stem.stem(t) for t in text] text = [word_lemmitize.lemmatize(t) for t in text] return " ".join(text)
3. 重新应用函数
处理完空值或者修改函数后,再执行列应用代码:
twitter_data['clean'] = twitter_data['comment'].apply(data_clean_pipeline)
这样应该就能顺利运行啦!把重复初始化的对象移到函数外,还能大幅提升整列数据的处理效率哦~
备注:内容来源于stack exchange,提问作者a_mittal
相关产品推荐
相关产品推荐

