运行SEO关键词聚类脚本时遇community_detection参数TypeError错误
解决community_detection()参数错误问题
问题原因
你调用的util.community_detection函数并不支持init_max_size这个参数,因此抛出类型错误。该参数要么是你记错了名称,要么是在当前使用的sentence-transformers版本中从未存在或已被移除。
解决方案
直接移除init_max_size=len(corpus_embeddings)这个参数即可,该函数的标准参数仅包含corpus_embeddings、min_community_size、threshold、batch_size和show_progress_bar。
修改后的关键代码行:
clusters = util.community_detection(corpus_embeddings, min_community_size=min_cluster_size, threshold=cluster_accuracy)
额外说明
如果你的需求是控制聚类的初始簇规模,可以通过调整以下参数实现类似效果:
- 降低
threshold:会让更多关键词被归为同一簇,簇整体规模更大 - 提高
min_community_size:会过滤掉过小的簇,只保留规模达标的簇
修改后的完整循环代码:
while cluster: corpus_sentences = list(corpus_set) check_len = len(corpus_sentences) corpus_embeddings = model.encode(corpus_sentences, batch_size=256, show_progress_bar=True, convert_to_tensor=True) # 移除不支持的init_max_size参数 clusters = util.community_detection(corpus_embeddings, min_community_size=min_cluster_size, threshold=cluster_accuracy) for keyword, cluster in enumerate(clusters): print("\nCluster {}, #{} Elements ".format(keyword + 1, len(cluster))) for sentence_id in cluster[0:]: print("\t", corpus_sentences[sentence_id]) corpus_sentences_list.append(corpus_sentences[sentence_id]) cluster_name_list.append("Cluster {}, #{} Elements ".format(keyword + 1, len(cluster))) df_new = pd.DataFrame(None) df_new['Cluster Name'] = cluster_name_list df_new["Keyword"] = corpus_sentences_list df_all.append(df_new) have = set(df_new["Keyword"]) corpus_set = corpus_set_all - have remaining = len(corpus_set) print("Total Unclustered Keywords: ", remaining) if check_len == remaining: break
内容的提问来源于stack exchange,提问作者sherif elshazly
相关产品推荐
相关产品推荐

