如何在for循环中正确显示pyLDAvis输出结果
问题原因
pyLDAvis.show()为阻塞式调用:方法启动本地HTTP服务后会持续占用当前线程,直到服务被手动终止才会执行后续代码,循环在第一次调用后就会卡住,不会进入后续类别的处理流程。- pyLDAvis 2.1.2版本默认启动的服务不会自动释放端口和进程资源,手动终止主程序时关联的子服务进程未被正常回收,就会抛出“process cannot be found”类报错。
解决方案
最稳定的方式是放弃循环内直接调用show()弹页面,改为将每个类别的可视化结果保存为独立HTML文件,所有类别处理完成后再逐个打开查看,不会出现阻塞、端口冲突问题。
修改后的代码如下:
import os # 新建文件夹存储所有可视化结果 os.makedirs("lda_vis_results", exist_ok=True) def lda_vis(text, category_name, nb_of_topics): tfidf_vectorizer = TfidfVectorizer(max_df=0.95, min_df=2, max_features=1000) tfidf = tfidf_vectorizer.fit_transform(text) lda = LatentDirichletAllocation(n_components=nb_of_topics, max_iter=5, learning_method='online', learning_offset=10, random_state=0) lda.fit(tfidf) visualisation = pyLDAvis.sklearn.prepare(lda, tfidf, tfidf_vectorizer) # 用类别名命名文件,避免结果覆盖 save_path = f"lda_vis_results/category_{category_name}_lda.html" pyLDAvis.save_html(visualisation, save_path) print(f"类别{category_name}的LDA可视化结果已保存至:{save_path}") for category in data['category'].unique(): df_by_category = data.loc[data['category'] == category] lda_vis(df_by_category['tokenized_sentence'], category_name=category, nb_of_topics=4)
如果确实需要循环执行时自动打开浏览器页面,可以给show()传入自定义端口避免冲突,该方案稳定性较差,旧版本pyLDAvis容易出现资源泄漏,仅作参考:
def lda_vis(text, nb_of_topics, port): tfidf_vectorizer = TfidfVectorizer(max_df=0.95, min_df=2, max_features=1000) tfidf = tfidf_vectorizer.fit_transform(text) lda = LatentDirichletAllocation(n_components=nb_of_topics, max_iter=5, learning_method='online', learning_offset=10, random_state=0) lda.fit(tfidf) visualisation = pyLDAvis.sklearn.prepare(lda, tfidf, tfidf_vectorizer) # 每次调用使用不同端口,避免冲突 pyLDAvis.show(visualisation, port=port, open_browser=True, local=True) # 从8000端口开始,每个类别分配独立端口 for idx, category in enumerate(data['category'].unique()): df_by_category = data.loc[data['category'] == category] lda_vis(df_by_category['tokenized_sentence'], nb_of_topics=4, port=8000+idx)
注意:使用该方案时,每打开一个可视化页面后,需要手动关闭对应运行的服务进程,循环才会进入下一轮处理。
内容的提问来源于stack exchange,提问作者Perrupi
相关产品推荐
相关产品推荐

