Lucene 7.x中CustomScoreQuery使用及getTermVector返回Null问题求助
嘿,我刚看完你的问题,正好之前也踩过Lucene词向量的坑!你遇到的getTermVector返回null的问题,核心原因其实很明确——Lucene的TextField默认是不存储词向量(Term Vectors)的,所以你就算索引了文档,也没法直接通过getTermVector拿到数据。下面给你一步步解决的办法:
1. 先确认问题根源
官方Demo里的IndexFiles是用new TextField("contents", reader)来创建内容字段的,而TextField的默认行为是:只做分词索引,不存储字段内容,也不存储词向量。所以你后续在自定义评分器里调用getTermVector自然会返回null。
2. 调整索引流程:开启词向量存储
要解决这个问题,你需要修改IndexFiles里创建contents字段的逻辑,手动指定字段类型并开启词向量存储。具体步骤如下:
在IndexFiles的indexDocs方法里,替换原来创建TextField的代码,改成自定义FieldType:
// 替换原来的 TextField 创建逻辑 FieldType contentFieldType = new FieldType(TextField.TYPE_NOT_STORED); contentFieldType.setStoreTermVectors(true); // 开启词向量存储 contentFieldType.setStoreTermVectorPositions(true); // 如果需要位置信息可选开启 contentFieldType.setStoreTermVectorOffsets(true); // 如果需要偏移量可选开启 contentFieldType.freeze(); // 冻结字段类型,防止后续修改 // 创建字段时使用这个自定义类型 Document doc = new Document(); doc.add(new Field("contents", reader, contentFieldType));
修改完成后,重新执行索引构建,然后用Luke工具查看索引:找到contents字段的属性,应该能看到Term Vectors已经标记为Yes了。这时候再运行你的CountingQuery,getTermVector就不会返回null了。
3. 替代方案:无需词向量获取文档词项信息(可选)
如果你因为某些原因不想重新构建索引,也可以通过LeafReader直接访问倒排表来获取单个文档的词项信息,不过这种方式相对复杂一点。示例代码如下(在你的customScore方法里修改):
public float customScore(int doc, float subQueryScore, float valSrcScores[]) throws IOException { LeafReader reader = context.reader(); // 获取contents字段的倒排表迭代器 Terms terms = reader.terms("contents"); if (terms != null) { TermsEnum termsEnum = terms.iterator(); PostingsEnum postingsEnum = null; BytesRef term; while ((term = termsEnum.next()) != null) { postingsEnum = termsEnum.postings(postingsEnum, PostingsEnum.FREQS); // 查找当前doc的词频 if (postingsEnum.advance(doc) == doc) { int freq = postingsEnum.freq(); // 这里可以处理词频逻辑 System.out.println("Term: " + term.utf8ToString() + ", Freq: " + freq); } } } return 1.0f; }
不过这种方式是遍历字段的所有词项,再检查当前文档是否包含该词项,性能上不如直接用词向量高效,所以如果可以重新构建索引的话,还是优先选择第一种方案。
内容的提问来源于stack exchange,提问作者S.Wagner

