使用Pandas生成词云时遇RecursionError,求优化方案
词云生成RecursionError优化方案
我正在可视化Kaggle数据集「Women In Headline: bias」,尝试用Pandas生成词云时遇到递归深度超出错误,原代码及报错如下:
原代码
def WordCloud (): data = Train_data[Train_data['bias'] == 5] text = " ".join(data["headline_no_site"]) text = "".join(_ for _ in text if _ not in punctuation) text = [t.lower() for t in text.split() if t.lower() not in Woman_word and t.lower() not in stpwrds and not t.isdigit()] #stopwords = set(STOPWORDS) wordcloud = WordCloud().generate(text)
报错信息
RecursionError: maximum recursion depth exceeded while calling a Python object
核心问题与优化方案
核心原因
错误根源是传给WordCloud.generate()的参数类型错误——你传入的是单词列表,但该方法要求输入字符串,这才触发了递归异常,和递归深度限制本身无关。
具体优化步骤
- 修正输入参数类型:将处理后的单词列表重新拼接为字符串,再传入
generate方法。 - 优化停用词查找效率:把停用词列表转为集合,集合的查找速度远快于列表,能大幅减少预处理的时间和内存占用。
- 简化标点过滤逻辑:用
str.translate()替代生成器过滤标点,效率更高。 - 降低词云计算压力:通过设置
max_words、调整图像分辨率等参数,减少内存消耗。
修正后的完整代码
import string from wordcloud import WordCloud def generate_wordcloud(): # 转换停用词为集合,提升查找效率 woman_word_set = set(Woman_word) stpwrds_set = set(stpwrds) # 筛选目标数据 data = Train_data[Train_data['bias'] == 5] # 拼接所有标题文本 text = " ".join(data["headline_no_site"]) # 快速过滤标点符号 translator = str.maketrans('', '', string.punctuation) text = text.translate(translator) # 预处理单词:转小写、过滤停用词与数字 processed_words = [t.lower() for t in text.split() if t.lower() not in woman_word_set and t.lower() not in stpwrds_set and not t.isdigit()] # 重新拼接为字符串(关键:符合WordCloud的输入要求) processed_text = " ".join(processed_words) # 生成词云,设置参数降低内存负载 wordcloud = WordCloud(max_words=200, width=800, height=600).generate(processed_text) return wordcloud
额外建议
如果数据集过大导致Colab崩溃,可以分批次处理文本:
- 先取小批量数据测试(比如
data = Train_data[Train_data['bias'] == 5].head(1000)),确认代码正常后再逐步扩大数据量。 - 用
pandas.read_csv的chunksize参数分块读取数据集,避免一次性加载全部数据到内存。
内容的提问来源于stack exchange,提问作者Willowinthewind
相关产品推荐
相关产品推荐

