You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Apache Lucene的StandardAnalyzer无法移除停用词

Lucene 10.0.0 StandardAnalyzer 默认不过滤停用词的问题

我用以下代码尝试移除字符串中的停用词,但未生效:

package com.example;

import java.io.IOException;
import java.util.ArrayList;
import java.util.List;

import org.apache.lucene.analysis.Analyzer;
import org.apache.lucene.analysis.TokenStream;
import org.apache.lucene.analysis.standard.StandardAnalyzer;
import org.apache.lucene.analysis.tokenattributes.CharTermAttribute;

public class Main {
    public static void main(String[] args) throws IOException {
        String text = "The quick brown fox jumps over the lazy dog";

        Analyzer analyzer = new StandardAnalyzer();
        TokenStream tokenStream = analyzer.tokenStream("field", text);
        CharTermAttribute charTermAttr = tokenStream.addAttribute(CharTermAttribute.class);

        tokenStream.reset();
        List<String> tokens = new ArrayList<>();
        while (tokenStream.incrementToken()) {
            tokens.add(charTermAttr.toString());
        }
        tokenStream.end();

        System.out.println("Tokens: " + tokens);
    }
}

对应的pom.xml配置:

<?xml version="1.0" encoding="UTF-8"?>
<project xmlns="http://maven.apache.org/POM/4.0.0"
    xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
    xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
    <modelVersion>4.0.0</modelVersion>

    <groupId>com.example</groupId>
    <artifactId>demo</artifactId>
    <version>1.0-SNAPSHOT</version>

    <properties>
        <maven.compiler.source>21</maven.compiler.source>
        <maven.compiler.target>21</maven.compiler.target>
    </properties>

    <dependencies>
        <dependency>
            <groupId>org.apache.lucene</groupId>
            <artifactId>lucene-core</artifactId>
            <version>10.0.0</version>
        </dependency>

        <dependency>
            <groupId>org.apache.lucene</groupId>
            <artifactId>lucene-queryparser</artifactId>
            <version>10.0.0</version>
        </dependency>


        <dependency>
            <groupId>org.apache.lucene</groupId>
            <artifactId>lucene-analysis-common</artifactId>
            <version>10.0.0</version>
        </dependency>
    </dependencies>
</project>

预期结果:Tokens: [quick, brown, fox, jumps, lazy, dog]
实际结果:Tokens: [the, quick, brown, fox, jumps, over, the, lazy, dog]

使用Lucene 10.0.0、Java 21,根据官方文档,StandardAnalyzer默认构造器应该构建无停用词的分析器,但实际没生效,请问原因是什么?


原因及解决方案

核心原因

Lucene 10.0.0的StandardAnalyzer默认构造器确实不会过滤任何停用词,你预期的结果是过滤英文停用词(如the、over),这是因为混淆了版本行为差异:

  • Lucene 9.x及更早版本中,StandardAnalyzer默认加载英文停用词表
  • 从Lucene 10.0.0开始,官方修改了默认实现,默认构造器不再包含停用词过滤器,因此不会自动过滤任何停用词

解决方案

如果需要过滤英文停用词,可选择以下两种方式:

  1. 使用自带英文停用词表构建分析器
    直接传入Lucene提供的英文停用词集合:
// 导入相关类
import org.apache.lucene.analysis.en.EnglishAnalyzer;

// 修改Analyzer初始化代码
Analyzer analyzer = new StandardAnalyzer(EnglishAnalyzer.ENGLISH_STOP_WORDS_SET);
  1. 自定义停用词集合
    如果需要自定义停用词,手动创建集合传入即可:
import java.util.HashSet;
import java.util.Set;

// 自定义停用词
Set<String> stopWords = new HashSet<>();
stopWords.add("the");
stopWords.add("over");
// 添加更多停用词...

Analyzer analyzer = new StandardAnalyzer(stopWords);

修改后运行代码,即可得到预期的过滤结果。


内容的提问来源于stack exchange,提问作者tle130475c

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 05:12:47