Pylucence 9.4.1无法检索已索引文档中存在词汇的问题求助
Pylucence 9.4.1索引后部分词汇无法检索的问题解决
问题描述
使用Pylucence 9.4.1对文档进行索引时,出现部分文档中已存在的词汇(例如baby)无法被检索到的异常。
索引代码
import os import pandas as pd from org.apache.lucene.document import Document, Field, FieldType from org.apache.lucene.index import IndexOptions, IndexWriter, IndexWriterConfig from org.apache.lucene.store import FSDirectory, File from org.apache.lucene.analysis.en import EnglishAnalyzer filepath = os.getcwd() + '/' + 'wiki_movie_plots_deduped.csv' def indexDocument(title, year, plot): ft = FieldType() ft.setIndexOptions(IndexOptions.DOCS_AND_FREQS_AND_POSITIONS_AND_OFFSETS); ft.setStored(True) ft.setTokenized(True) ft.setStoreTermVectors(True) ft.setStoreTermVectorOffsets(True) ft.setStoreTermVectorPositions(True) doc = Document() doc.add(Field("Title", title, ft)) doc.add(Field("Plot", plot, ft)) writer.addDocument(doc) def CloseWriter(): writer.close() def makeInvertedIndex(file_path): df = pd.read_csv(file_path) print(df.columns) docid = 0 for i in df.index: print(docid, '-', df['Title'][i]) indexDocument(df['Title'][i], df['Release Year'][i], df['Plot'][i]) docid += 1 indexPath = File('index/').toPath() indexDir = FSDirectory.open(indexPath) writerConfig = IndexWriterConfig(EnglishAnalyzer()) writer = IndexWriter(indexDir, writerConfig) inverted = makeInvertedIndex(filepath) CloseWriter()
检索代码
from org.apache.lucene.search import IndexSearcher, BM25Similarity from org.apache.lucene.index import DirectoryReader from org.apache.lucene.queryparser.classic import QueryParser from org.apache.lucene.store import FSDirectory, File from org.apache.lucene.analysis.standard import StandardAnalyzer keyword = 'baby' fieldname = 'Title' result = list() indexPath = File('index/').toPath() directory = FSDirectory.open(indexPath) analyzer = StandardAnalyzer() reader = DirectoryReader.open(directory) searcher = IndexSearcher(DirectoryReader.open(directory)) query = QueryParser(fieldname, analyzer).parse(keyword) print('query', query) numdocs = searcher.count(query) print("#-docs:", numdocs) searcher.setSimilarity(BM25Similarity(1.2,0.75)) scoreDocs = searcher.search(query, 1000).scoreDocs # it returns TopDocs object containing scoreDocs and totalHits # scoreDoc object contains docId and score print('total hit:', searcher.search(query, 100).totalHits) print("%s total matching documents" % (len(scoreDocs)))
问题分析
核心问题是索引与检索阶段使用的文本分析器(Analyzer)不一致:
- 索引阶段使用
EnglishAnalyzer(),它包含英文特有的处理逻辑:词干提取(如将Babies转换为baby)、停用词过滤(移除the、a等常见词)、小写转换。 - 检索阶段使用
StandardAnalyzer(),仅做基础分词和小写转换,缺少词干提取和停用词过滤的逻辑。
这种不一致会导致查询词经过处理后,与索引中存储的词项无法匹配,最终出现检索不到的情况。
解决方案
1. 统一索引与检索的Analyzer
将检索代码中的StandardAnalyzer()替换为EnglishAnalyzer(),确保两端的文本处理逻辑完全一致:
# 替换原有的StandardAnalyzer from org.apache.lucene.analysis.en import EnglishAnalyzer analyzer = EnglishAnalyzer()
2. 验证索引词项(调试用)
若问题仍存在,可通过以下代码查看索引中实际存储的词项,对比查询词处理后的结果,定位匹配问题:
from org.apache.lucene.index import MultiFields reader = DirectoryReader.open(directory) terms = MultiFields.getTerms(reader, "Title") if terms is not None: terms_enum = terms.iterator() term = terms_enum.next() while term is not None: print(term.utf8ToString()) term = terms_enum.next() reader.close()
3. 确认FieldType配置
确保FieldType已开启分词(你的代码中ft.setTokenized(True)已正确设置,此步骤为验证用),分词是Analyzer生效的前提。
内容的提问来源于stack exchange,提问作者NFStudio
相关产品推荐
相关产品推荐

