You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Lucene 10.0.0读取老旧Apache Lucene索引失败求助

读取旧Lucene索引失败的解决方案

问题背景

作为Lucene和Java新手,尝试读取一个8-10年历史的旧索引,使用Lucene 10.0.0和JDK 23.0时触发版本兼容错误,尝试过Lucene 3.0.3等版本仍无法读取,仅需提取索引数据后用最新版本重建索引。

系统环境

  • Windows 10
  • Lucene 10.0.0
  • JDK 23.0

索引目录文件

名称大小
_0.cfx47,942 KB
_s.cfs178,687 KB
segments.gen1 KB
segments_21 KB

测试代码

import java.util.List;
import java.io.IOException;
import java.nio.file.Paths;

import org.apache.lucene.document.Document;
import org.apache.lucene.index.DirectoryReader;
import org.apache.lucene.index.IndexReader;
import org.apache.lucene.index.StoredFields;
import org.apache.lucene.index.IndexableField;
import org.apache.lucene.search.IndexSearcher;
import org.apache.lucene.search.Query;
import org.apache.lucene.search.ScoreDoc;
import org.apache.lucene.search.TopDocs;
import org.apache.lucene.store.FSDirectory;
import org.apache.lucene.search.MatchAllDocsQuery;

public class FieldReader {

    private static final String INDEX_PATH = "indexedfiles";

    public static void main(String[] args) throws Exception {
        IndexReader reader = DirectoryReader.open(FSDirectory.open(Paths.get(INDEX_PATH)));
        //Index searcher
        IndexSearcher searcher = new IndexSearcher(reader);
        //This query will match with all documents in the index
        Query query = new MatchAllDocsQuery();
        //search the index
        TopDocs foundDocs = searcher.search(query, 10);
        // Returns a StoredFields reader for the stored fields of this index. 
        StoredFields storedFields = searcher.storedFields();

        //Let's print out the path of document files and fields
        for (ScoreDoc sd : foundDocs.scoreDocs) {
          Document doc = storedFields.document(sd.doc);
          System.out.println("Path : " + doc.get("path"));
          List<IndexableField> fields = doc.getFields();
            for (IndexableField field : fields) {
                System.out.println("Name : " + field.name() + ", Type : " + field.fieldType().toString());
            }
        }
    }
}

报错信息

Exception in thread "main" org.apache.lucene.index.IndexFormatTooOldException:
Format version is not supported
(resource BufferedChecksumIndexInput(MemorySegmentIndexInput(path="D:\IndiaLawLibrary\AHC\s_index\a\segments_2"))): -9 (needs to be between 1071082519 and 1071082519).
This version of Lucene only supports indexes created with release 9.0 and later.

解决步骤

1. 定位索引对应的Lucene版本

报错中的版本号-9对应Lucene 3.1.x版本(早期Lucene用负数标识segments格式版本,-9是3.1的专属版本号),3.0.3与3.1索引格式不兼容,必须使用Lucene 3.1.0或3.1.1读取。

2. 搭建兼容运行环境

  • JDK版本:Lucene 3.1.x最高适配JDK 1.7,建议使用JDK 1.7(避免高版本JDK的语法或API兼容问题)。
  • 依赖准备:下载Lucene 3.1.1核心jar包(lucene-core-3.1.1.jar),读取存储字段无需额外依赖。

3. 适配Lucene 3.1 API编写读取代码

Lucene 3.x与10.x API差异较大,修改后的读取代码如下:

import java.io.File;
import java.util.List;

import org.apache.lucene.document.Document;
import org.apache.lucene.document.Field;
import org.apache.lucene.index.IndexReader;
import org.apache.lucene.search.IndexSearcher;
import org.apache.lucene.search.MatchAllDocsQuery;
import org.apache.lucene.search.ScoreDoc;
import org.apache.lucene.search.TopDocs;
import org.apache.lucene.store.FSDirectory;

public class OldIndexReader {
    private static final String OLD_INDEX_PATH = "D:\\IndiaLawLibrary\\AHC\\s_index\\a";

    public static void main(String[] args) throws Exception {
        // 打开旧索引
        IndexReader reader = IndexReader.open(FSDirectory.open(new File(OLD_INDEX_PATH)));
        IndexSearcher searcher = new IndexSearcher(reader);
        
        // 查询所有文档
        MatchAllDocsQuery query = new MatchAllDocsQuery();
        TopDocs topDocs = searcher.search(query, reader.maxDoc());
        
        // 遍历提取存储字段数据
        for (ScoreDoc sd : topDocs.scoreDocs) {
            Document doc = reader.document(sd.doc);
            System.out.println("Path : " + doc.get("path"));
            List<Field> fields = doc.getFields();
            for (Field field : fields) {
                System.out.println("Name : " + field.name() + ", Type : " + field.fieldType());
            }
        }
        
        reader.close();
    }
}

4. 迁移数据到新索引

  • 运行上述代码提取所有文档的存储字段后,使用Lucene 10.x的API创建新索引,将提取的数据写入。
  • 注意字段类型映射:旧版本的Field.Store.YES对应新版本的StoredField或StringField(根据是否需要索引调整)。

依赖获取

若使用Maven,可添加以下依赖坐标下载Lucene 3.1.1:

<dependency>
    <groupId>org.apache.lucene</groupId>
    <artifactId>lucene-core</artifactId>
    <version>3.1.1</version>
</dependency>

内容的提问来源于stack exchange,提问作者Prashant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 04:59:52