You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pylucence 9.4.1无法检索已索引文档中存在词汇的问题求助

Pylucence 9.4.1索引后部分词汇无法检索的问题解决

问题描述

使用Pylucence 9.4.1对文档进行索引时,出现部分文档中已存在的词汇(例如baby)无法被检索到的异常。

索引代码

import os
import pandas as pd
from org.apache.lucene.document import Document, Field, FieldType
from org.apache.lucene.index import IndexOptions, IndexWriter, IndexWriterConfig
from org.apache.lucene.store import FSDirectory, File
from org.apache.lucene.analysis.en import EnglishAnalyzer

filepath = os.getcwd() + '/' + 'wiki_movie_plots_deduped.csv'


def indexDocument(title, year, plot):
    ft = FieldType()
    ft.setIndexOptions(IndexOptions.DOCS_AND_FREQS_AND_POSITIONS_AND_OFFSETS);
    ft.setStored(True)
    ft.setTokenized(True)
    ft.setStoreTermVectors(True)
    ft.setStoreTermVectorOffsets(True)
    ft.setStoreTermVectorPositions(True)
    doc = Document()
    doc.add(Field("Title", title, ft))
    doc.add(Field("Plot", plot, ft))    
    writer.addDocument(doc)


def CloseWriter():
    writer.close()
    

def makeInvertedIndex(file_path):
    df = pd.read_csv(file_path)
    print(df.columns)
    docid = 0
    for i in df.index:
        print(docid, '-', df['Title'][i])
        indexDocument(df['Title'][i], df['Release Year'][i], df['Plot'][i])
        docid += 1
  

indexPath = File('index/').toPath()
indexDir = FSDirectory.open(indexPath)
writerConfig = IndexWriterConfig(EnglishAnalyzer())
writer = IndexWriter(indexDir, writerConfig)
inverted = makeInvertedIndex(filepath)

CloseWriter()

检索代码

from org.apache.lucene.search import IndexSearcher, BM25Similarity
from org.apache.lucene.index import DirectoryReader
from org.apache.lucene.queryparser.classic import QueryParser
from org.apache.lucene.store import FSDirectory, File
from org.apache.lucene.analysis.standard import StandardAnalyzer

keyword = 'baby'
fieldname = 'Title'
result = list()

indexPath = File('index/').toPath()
directory = FSDirectory.open(indexPath)

analyzer = StandardAnalyzer()
reader = DirectoryReader.open(directory)
searcher = IndexSearcher(DirectoryReader.open(directory))
query = QueryParser(fieldname, analyzer).parse(keyword)
print('query', query)
numdocs = searcher.count(query)
print("#-docs:", numdocs)
    

searcher.setSimilarity(BM25Similarity(1.2,0.75))
scoreDocs = searcher.search(query, 1000).scoreDocs # it returns TopDocs object containing scoreDocs and totalHits
# scoreDoc object contains docId and score
print('total hit:', searcher.search(query, 100).totalHits)
print("%s total matching documents" % (len(scoreDocs)))

问题分析

核心问题是索引与检索阶段使用的文本分析器(Analyzer)不一致:

  • 索引阶段使用EnglishAnalyzer(),它包含英文特有的处理逻辑:词干提取(如将Babies转换为baby)、停用词过滤(移除the、a等常见词)、小写转换。
  • 检索阶段使用StandardAnalyzer(),仅做基础分词和小写转换,缺少词干提取和停用词过滤的逻辑。

这种不一致会导致查询词经过处理后,与索引中存储的词项无法匹配,最终出现检索不到的情况。

解决方案

1. 统一索引与检索的Analyzer

将检索代码中的StandardAnalyzer()替换为EnglishAnalyzer(),确保两端的文本处理逻辑完全一致:

# 替换原有的StandardAnalyzer
from org.apache.lucene.analysis.en import EnglishAnalyzer
analyzer = EnglishAnalyzer()

2. 验证索引词项(调试用)

若问题仍存在,可通过以下代码查看索引中实际存储的词项,对比查询词处理后的结果,定位匹配问题:

from org.apache.lucene.index import MultiFields

reader = DirectoryReader.open(directory)
terms = MultiFields.getTerms(reader, "Title")
if terms is not None:
    terms_enum = terms.iterator()
    term = terms_enum.next()
    while term is not None:
        print(term.utf8ToString())
        term = terms_enum.next()
reader.close()

3. 确认FieldType配置

确保FieldType已开启分词(你的代码中ft.setTokenized(True)已正确设置,此步骤为验证用),分词是Analyzer生效的前提。

内容的提问来源于stack exchange,提问作者NFStudio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 11:05:22