You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Solr中获取排除通用词的Top关键词(多语言数据集)

针对Solr多语言数据集提取Top10关键词的解决方案

一、优化停用词过滤(替代手动罗列facet.excludeTerms)

针对多语言场景,无需在查询阶段手动枚举通用词,可从索引层面配置多语言停用词过滤:

  1. 准备多语言停用词集合:将英、法、西语等通用停用词合并为multilingual_stopwords.txt,放置到Solr核心的conf目录下。
  2. 修改字段类型配置(schema.xml或managed-schema):
    <fieldType name="text_multilingual" class="solr.TextField" positionIncrementGap="100">
      <analyzer type="index">
        <tokenizer class="solr.StandardTokenizerFactory"/>
        <filter class="solr.LowerCaseFilterFactory"/>
        <filter class="solr.StopFilterFactory" words="multilingual_stopwords.txt" ignoreCase="true"/>
        <!-- 可选:过滤1-2字符的短词,进一步剔除无意义词汇 -->
        <filter class="solr.LengthFilterFactory" min="3" max="100"/>
      </analyzer>
      <analyzer type="query">
        <tokenizer class="solr.StandardTokenizerFactory"/>
        <filter class="solr.LowerCaseFilterFactory"/>
        <filter class="solr.StopFilterFactory" words="multilingual_stopwords.txt" ignoreCase="true"/>
      </analyzer>
    </fieldType>
    
  3. 将content字段的类型改为text_multilingual,重新索引数据后,直接用简化后的facet查询即可自动排除通用词:
    http://localhost:8983/solr/<my_core>/select?facet=true&facet.field=content&facet.limit=10&facet.minCount=1&q=*:*&rows=0
    
    (rows=0仅返回facet结果,无需返回文档;q=*:*替代原content:(%2A)更简洁)

二、修复tf-idf函数的使用问题

你之前的tf-idf查询存在两个核心问题:

  1. tf值为0的原因:tf(content,'covid')是获取单篇返回文档的词频,若查询返回的文档不含目标词,tf值自然为0。全局词频统计不能依赖文档级的tf函数,需结合facet或词向量组件。
  2. idf结果异常的原因:idf公式为log(总文档数/包含该词的文档数),若and的idf值高于covid,说明你的数据集中包含and的文档数更少——本质是索引阶段未过滤停用词,导致通用词的分布不符合预期。

正确利用tf-idf筛选关键词的步骤:

  1. 获取总文档数:
    http://localhost:8983/solr/<my_core>/select?q=*:*&rows=0&wt=json
    
    从response.numFound中得到总文档数N。
  2. 获取带文档频率(df)的词统计:
    先在schema中给content字段开启文档频率返回:
    <field name="content" type="text_multilingual" indexed="true" stored="true" facet="true" facet.returnDF="true"/>
    
    再执行查询:
    http://localhost:8983/solr/<my_core>/select?facet=true&facet.field=content&facet.limit=-1&facet.minCount=1&facet.method=enum&q=*:*&rows=0&wt=json
    
    • facet.limit=-1返回所有词的统计结果
    • facet.method=enum确保返回每个词的文档频率(df)
  3. 计算tf-idf得分:tf_idf = 词总出现次数(tf) * log(N/包含该词的文档数(df)),排序后取Top10即可。

三、替代方案:使用TermVectorComponent提取关键词

启用词向量组件可直接获取全局词频与文档频率:

  1. 修改schema开启词向量:
    <field name="content" type="text_multilingual" indexed="true" stored="true" termVectors="true" termPositions="false" termOffsets="false" termPayloads="false"/>
    
  2. 重新索引后,执行全局词向量查询:
    http://localhost:8983/solr/<my_core>/tv?q=*:*&tv.fl=content&tv.global=true&wt=json
    
    从返回结果中提取每个词的tf和df,计算tf-idf后排序取Top10。

内容的提问来源于stack exchange,提问作者vanessa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 19:31:09