You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行SEO关键词聚类脚本时遇community_detection参数TypeError错误

解决community_detection()参数错误问题

问题原因

你调用的util.community_detection函数并不支持init_max_size这个参数,因此抛出类型错误。该参数要么是你记错了名称,要么是在当前使用的sentence-transformers版本中从未存在或已被移除。

解决方案

直接移除init_max_size=len(corpus_embeddings)这个参数即可,该函数的标准参数仅包含corpus_embeddings、min_community_size、threshold、batch_size和show_progress_bar。

修改后的关键代码行:

clusters = util.community_detection(corpus_embeddings, min_community_size=min_cluster_size, threshold=cluster_accuracy)

额外说明

如果你的需求是控制聚类的初始簇规模,可以通过调整以下参数实现类似效果:

  • 降低threshold:会让更多关键词被归为同一簇,簇整体规模更大
  • 提高min_community_size:会过滤掉过小的簇,只保留规模达标的簇

修改后的完整循环代码:

while cluster:
    corpus_sentences = list(corpus_set)
    check_len = len(corpus_sentences)

    corpus_embeddings = model.encode(corpus_sentences, batch_size=256, show_progress_bar=True, convert_to_tensor=True)
    # 移除不支持的init_max_size参数
    clusters = util.community_detection(corpus_embeddings, min_community_size=min_cluster_size, threshold=cluster_accuracy)

    for keyword, cluster in enumerate(clusters):
        print("\nCluster {}, #{} Elements ".format(keyword + 1, len(cluster)))

        for sentence_id in cluster[0:]:
            print("\t", corpus_sentences[sentence_id])
            corpus_sentences_list.append(corpus_sentences[sentence_id])
            cluster_name_list.append("Cluster {}, #{} Elements ".format(keyword + 1, len(cluster)))

    df_new = pd.DataFrame(None)
    df_new['Cluster Name'] = cluster_name_list
    df_new["Keyword"] = corpus_sentences_list

    df_all.append(df_new)
    have = set(df_new["Keyword"])

    corpus_set = corpus_set_all - have
    remaining = len(corpus_set)
    print("Total Unclustered Keywords: ", remaining)
    if check_len == remaining:
        break

内容的提问来源于stack exchange,提问作者sherif elshazly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 07:25:17