如何在Solr中获取排除通用词的Top关键词(多语言数据集)
针对Solr多语言数据集提取Top10关键词的解决方案
一、优化停用词过滤(替代手动罗列facet.excludeTerms)
针对多语言场景,无需在查询阶段手动枚举通用词,可从索引层面配置多语言停用词过滤:
- 准备多语言停用词集合:将英、法、西语等通用停用词合并为
multilingual_stopwords.txt,放置到Solr核心的conf目录下。 - 修改字段类型配置(
schema.xml或managed-schema):<fieldType name="text_multilingual" class="solr.TextField" positionIncrementGap="100"> <analyzer type="index"> <tokenizer class="solr.StandardTokenizerFactory"/> <filter class="solr.LowerCaseFilterFactory"/> <filter class="solr.StopFilterFactory" words="multilingual_stopwords.txt" ignoreCase="true"/> <!-- 可选:过滤1-2字符的短词,进一步剔除无意义词汇 --> <filter class="solr.LengthFilterFactory" min="3" max="100"/> </analyzer> <analyzer type="query"> <tokenizer class="solr.StandardTokenizerFactory"/> <filter class="solr.LowerCaseFilterFactory"/> <filter class="solr.StopFilterFactory" words="multilingual_stopwords.txt" ignoreCase="true"/> </analyzer> </fieldType> - 将
content字段的类型改为text_multilingual,重新索引数据后,直接用简化后的facet查询即可自动排除通用词:
(http://localhost:8983/solr/<my_core>/select?facet=true&facet.field=content&facet.limit=10&facet.minCount=1&q=*:*&rows=0rows=0仅返回facet结果,无需返回文档;q=*:*替代原content:(%2A)更简洁)
二、修复tf-idf函数的使用问题
你之前的tf-idf查询存在两个核心问题:
- tf值为0的原因:
tf(content,'covid')是获取单篇返回文档的词频,若查询返回的文档不含目标词,tf值自然为0。全局词频统计不能依赖文档级的tf函数,需结合facet或词向量组件。 - idf结果异常的原因:idf公式为
log(总文档数/包含该词的文档数),若and的idf值高于covid,说明你的数据集中包含and的文档数更少——本质是索引阶段未过滤停用词,导致通用词的分布不符合预期。
正确利用tf-idf筛选关键词的步骤:
- 获取总文档数:
从http://localhost:8983/solr/<my_core>/select?q=*:*&rows=0&wt=jsonresponse.numFound中得到总文档数N。 - 获取带文档频率(df)的词统计:
先在schema中给content字段开启文档频率返回:
再执行查询:<field name="content" type="text_multilingual" indexed="true" stored="true" facet="true" facet.returnDF="true"/>http://localhost:8983/solr/<my_core>/select?facet=true&facet.field=content&facet.limit=-1&facet.minCount=1&facet.method=enum&q=*:*&rows=0&wt=jsonfacet.limit=-1返回所有词的统计结果facet.method=enum确保返回每个词的文档频率(df)
- 计算tf-idf得分:
tf_idf = 词总出现次数(tf) * log(N/包含该词的文档数(df)),排序后取Top10即可。
三、替代方案:使用TermVectorComponent提取关键词
启用词向量组件可直接获取全局词频与文档频率:
- 修改schema开启词向量:
<field name="content" type="text_multilingual" indexed="true" stored="true" termVectors="true" termPositions="false" termOffsets="false" termPayloads="false"/> - 重新索引后,执行全局词向量查询:
从返回结果中提取每个词的tf和df,计算tf-idf后排序取Top10。http://localhost:8983/solr/<my_core>/tv?q=*:*&tv.fl=content&tv.global=true&wt=json
内容的提问来源于stack exchange,提问作者vanessa
相关产品推荐
相关产品推荐

