You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用SortField.FIELD_SCORE排序时Apache Lucene返回NaN分数问题

Lucene使用SortField.FIELD_SCORE排序时ScoreDoc.score返回NaN的问题

问题描述

想要将Apache Lucene搜索结果按相关性排序,但使用SortField.FIELD_SCORE进行排序时,返回文档的分数始终为NaN;省略排序参数时,搜索正常且文档包含有效分数。

使用的依赖版本:Maven仓库中的lucene-core 9.6.0、lucene-analyzers-common 8.11.2。

可复现问题的极简代码:

public class LuceneSearch {
    public static void main(String[] args) {
        try {
            Directory directory = new ByteBuffersDirectory();
            
            try (IndexWriter indexWriter = new IndexWriter(directory, new IndexWriterConfig(new SimpleAnalyzer()))) {
                indexWriter.addDocument(createDocument("a very simple example"));
                indexWriter.addDocument(createDocument("another example"));
                indexWriter.addDocument(createDocument("hello world"));
            }

            IndexReader indexReader = DirectoryReader.open(directory);
            IndexSearcher indexSearcher = new IndexSearcher(indexReader);

            Query query = new TermQuery(new Term("value", "hello"));
            Sort sort = new Sort(SortField.FIELD_SCORE); // <<<< 引发问题的代码
            TopDocs topDocs = indexSearcher.search(query, 10, sort);
            for (ScoreDoc scoreDoc : topDocs.scoreDocs) {
                System.out.println(scoreDoc.doc + " : " + scoreDoc.score);
            }

            indexReader.close();
            directory.close();
        } catch (IOException e) {
            e.printStackTrace();
        }
    }

    private static Document createDocument(String value) {
        Document document = new Document();
        document.add(new TextField("value", value, Field.Store.NO));
        return document;
    }
}

运行结果:

  • 设置排序参数时输出:2 : NaN
  • 不设置排序参数时输出:2 : 0.49662238

编辑补充:实际返回的ScoreDoc是FieldDoc实例,它的fields属性包含排序时的有效分数,测试发现有无排序参数时实际分数一致,排序逻辑正常。

问题原因与解决方案

这不是Lucene的Bug,而是显式指定排序规则时的设计行为:

  • 显式使用SortField.FIELD_SCORE排序时,Lucene会将排序用的分数存入FieldDoc的fields数组,不会填充ScoreDoc的score字段,因此该字段显示为NaN。
  • 不指定排序参数时,Lucene默认按相关性排序,此时会直接填充ScoreDoc的score字段。

获取显式排序时的分数

可以将ScoreDoc强转为FieldDoc,从fields属性中取出有效分数:

for (ScoreDoc scoreDoc : topDocs.scoreDocs) {
    if (scoreDoc instanceof FieldDoc fieldDoc) {
        // fields[0]对应SortField.FIELD_SCORE的排序值
        float actualScore = (float) fieldDoc.fields[0];
        System.out.println(scoreDoc.doc + " : " + actualScore);
    }
}

额外建议

你当前使用的lucene-core(9.6.0)和lucene-analyzers-common(8.11.2)版本不一致,虽然当前问题与版本无关,但建议保持Lucene各组件版本统一,避免潜在的兼容性问题。

内容的提问来源于stack exchange,提问作者Franz Deschler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 07:53:24