使用Lucene 10.0.0读取老旧Apache Lucene索引失败求助
读取旧Lucene索引失败的解决方案
问题背景
作为Lucene和Java新手,尝试读取一个8-10年历史的旧索引,使用Lucene 10.0.0和JDK 23.0时触发版本兼容错误,尝试过Lucene 3.0.3等版本仍无法读取,仅需提取索引数据后用最新版本重建索引。
系统环境
- Windows 10
- Lucene 10.0.0
- JDK 23.0
索引目录文件
| 名称 | 大小 |
|---|---|
| _0.cfx | 47,942 KB |
| _s.cfs | 178,687 KB |
| segments.gen | 1 KB |
| segments_2 | 1 KB |
测试代码
import java.util.List; import java.io.IOException; import java.nio.file.Paths; import org.apache.lucene.document.Document; import org.apache.lucene.index.DirectoryReader; import org.apache.lucene.index.IndexReader; import org.apache.lucene.index.StoredFields; import org.apache.lucene.index.IndexableField; import org.apache.lucene.search.IndexSearcher; import org.apache.lucene.search.Query; import org.apache.lucene.search.ScoreDoc; import org.apache.lucene.search.TopDocs; import org.apache.lucene.store.FSDirectory; import org.apache.lucene.search.MatchAllDocsQuery; public class FieldReader { private static final String INDEX_PATH = "indexedfiles"; public static void main(String[] args) throws Exception { IndexReader reader = DirectoryReader.open(FSDirectory.open(Paths.get(INDEX_PATH))); //Index searcher IndexSearcher searcher = new IndexSearcher(reader); //This query will match with all documents in the index Query query = new MatchAllDocsQuery(); //search the index TopDocs foundDocs = searcher.search(query, 10); // Returns a StoredFields reader for the stored fields of this index. StoredFields storedFields = searcher.storedFields(); //Let's print out the path of document files and fields for (ScoreDoc sd : foundDocs.scoreDocs) { Document doc = storedFields.document(sd.doc); System.out.println("Path : " + doc.get("path")); List<IndexableField> fields = doc.getFields(); for (IndexableField field : fields) { System.out.println("Name : " + field.name() + ", Type : " + field.fieldType().toString()); } } } }
报错信息
Exception in thread "main" org.apache.lucene.index.IndexFormatTooOldException:
Format version is not supported
(resource BufferedChecksumIndexInput(MemorySegmentIndexInput(path="D:\IndiaLawLibrary\AHC\s_index\a\segments_2"))): -9 (needs to be between 1071082519 and 1071082519).
This version of Lucene only supports indexes created with release 9.0 and later.
解决步骤
1. 定位索引对应的Lucene版本
报错中的版本号-9对应Lucene 3.1.x版本(早期Lucene用负数标识segments格式版本,-9是3.1的专属版本号),3.0.3与3.1索引格式不兼容,必须使用Lucene 3.1.0或3.1.1读取。
2. 搭建兼容运行环境
- JDK版本:Lucene 3.1.x最高适配JDK 1.7,建议使用JDK 1.7(避免高版本JDK的语法或API兼容问题)。
- 依赖准备:下载Lucene 3.1.1核心jar包(lucene-core-3.1.1.jar),读取存储字段无需额外依赖。
3. 适配Lucene 3.1 API编写读取代码
Lucene 3.x与10.x API差异较大,修改后的读取代码如下:
import java.io.File; import java.util.List; import org.apache.lucene.document.Document; import org.apache.lucene.document.Field; import org.apache.lucene.index.IndexReader; import org.apache.lucene.search.IndexSearcher; import org.apache.lucene.search.MatchAllDocsQuery; import org.apache.lucene.search.ScoreDoc; import org.apache.lucene.search.TopDocs; import org.apache.lucene.store.FSDirectory; public class OldIndexReader { private static final String OLD_INDEX_PATH = "D:\\IndiaLawLibrary\\AHC\\s_index\\a"; public static void main(String[] args) throws Exception { // 打开旧索引 IndexReader reader = IndexReader.open(FSDirectory.open(new File(OLD_INDEX_PATH))); IndexSearcher searcher = new IndexSearcher(reader); // 查询所有文档 MatchAllDocsQuery query = new MatchAllDocsQuery(); TopDocs topDocs = searcher.search(query, reader.maxDoc()); // 遍历提取存储字段数据 for (ScoreDoc sd : topDocs.scoreDocs) { Document doc = reader.document(sd.doc); System.out.println("Path : " + doc.get("path")); List<Field> fields = doc.getFields(); for (Field field : fields) { System.out.println("Name : " + field.name() + ", Type : " + field.fieldType()); } } reader.close(); } }
4. 迁移数据到新索引
- 运行上述代码提取所有文档的存储字段后,使用Lucene 10.x的API创建新索引,将提取的数据写入。
- 注意字段类型映射:旧版本的
Field.Store.YES对应新版本的StoredField或StringField(根据是否需要索引调整)。
依赖获取
若使用Maven,可添加以下依赖坐标下载Lucene 3.1.1:
<dependency> <groupId>org.apache.lucene</groupId> <artifactId>lucene-core</artifactId> <version>3.1.1</version> </dependency>
内容的提问来源于stack exchange,提问作者Prashant
相关产品推荐
相关产品推荐

