You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Spring Data Elasticsearch中创建并使用Ingest Pipeline处理PDF索引

Spring Boot集成Elasticsearch Ingest Pipeline实现PDF索引与检索

一、先确认Pipeline状态

既然你已经用Postman验证过Pipeline能正常工作,首先要确保这个Pipeline已经在Elasticsearch实例中存在——比如处理PDF常用的attachment处理器Pipeline,假设你的Pipeline名叫pdf-attachment-pipeline。如果需要在Spring Boot项目里动态创建Pipeline,也可以用客户端API来实现(后面会附代码示例)。

二、Spring Boot配置Elasticsearch客户端

以Spring Boot 3.x + Spring Data Elasticsearch 5.x为例,先完成客户端基础配置:

1. 依赖配置

在pom.xml中添加Spring Data Elasticsearch starter:

<dependency>
    <groupId>org.springframework.boot</groupId>
    <artifactId>spring-boot-starter-data-elasticsearch</artifactId>
</dependency>

2. 客户端配置类

import org.springframework.context.annotation.Configuration;
import org.springframework.data.elasticsearch.client.ClientConfiguration;
import org.springframework.data.elasticsearch.client.elc.ElasticsearchConfiguration;

@Configuration
public class ElasticsearchConfig extends ElasticsearchConfiguration {

    @Override
    public ClientConfiguration clientConfiguration() {
        return ClientConfiguration.builder()
                .connectedTo("localhost:9200") // 替换成你的ES地址
                // 有账号密码的话打开下面注释
                // .withBasicAuth("your-username", "your-password")
                .build();
    }
}

三、两种方式结合Pipeline处理PDF

方式1:直接上传二进制文件(无需实体类)

适合直接处理文件上传的场景,用ElasticsearchClient直接提交二进制数据并指定Pipeline:

import org.springframework.data.elasticsearch.client.elc.ElasticsearchClient;
import org.springframework.stereotype.Service;
import java.io.File;
import java.io.IOException;
import java.nio.file.Files;

@Service
public class PdfIndexService {

    private final ElasticsearchClient elasticsearchClient;

    public PdfIndexService(ElasticsearchClient elasticsearchClient) {
        this.elasticsearchClient = elasticsearchClient;
    }

    // 索引PDF文件
    public void indexPdfFile(File pdfFile) throws IOException {
        byte[] pdfBytes = Files.readAllBytes(pdfFile.toPath());
        
        // 提交索引请求,指定使用的Pipeline
        elasticsearchClient.index(i -> i
                .index("pdf-docs") // 替换成你的索引名
                .pipeline("pdf-attachment-pipeline") // 你的Pipeline名称
                .document(new PdfUploadDto(pdfBytes, pdfFile.getName()))
        );
    }

    // 封装PDF数据的简单DTO
    private static class PdfUploadDto {
        private byte[] data;
        private String fileName;

        public PdfUploadDto(byte[] data, String fileName) {
            this.data = data;
            this.fileName = fileName;
        }

        // getters
        public byte[] getData() { return data; }
        public String getFileName() { return fileName; }
    }

    // 可选:代码创建Pipeline(如果还没在ES里创建)
    public void createPdfPipeline() throws IOException {
        elasticsearchClient.putPipeline(p -> p
                .id("pdf-attachment-pipeline")
                .description("从PDF提取文本内容")
                .processors(pr -> pr
                        .attachment(a -> a
                                .field("data") // 对应DTO里的data字段
                                .targetField("attachment_content") // 提取后的文本存到这个字段
                                .properties(pf -> pf.add("content")) // 只提取文本内容,可按需添加其他属性
                        )
                        .remove(r -> r.field("data")) // 可选:提取完成后删除原始二进制数据,节省存储空间
                )
        );
    }
}

方式2:结合实体类注解使用Pipeline

如果你习惯用实体类映射ES文档,可以这么做:

1. 定义实体类

import org.springframework.data.annotation.Id;
import org.springframework.data.elasticsearch.annotations.Document;
import org.springframework.data.elasticsearch.annotations.Field;
import org.springframework.data.elasticsearch.annotations.FieldType;

@Document(indexName = "pdf-docs")
public class PdfDoc {

    @Id
    private String id;

    private String fileName;

    // 存储PDF二进制数据,由Pipeline处理
    @Field(type = FieldType.Binary)
    private byte[] data;

    // Pipeline提取后的文本内容,无需手动赋值,ES自动填充
    @Field(type = FieldType.Text, analyzer = "ik_max_word") // 中文建议用ik分词器,英文用standard
    private String attachmentContent;

    // 标准getter、setter
    public String getId() { return id; }
    public void setId(String id) { this.id = id; }
    public String getFileName() { return fileName; }
    public void setFileName(String fileName) { this.fileName = fileName; }
    public byte[] getData() { return data; }
    public void setData(byte[] data) { this.data = data; }
    public String getAttachmentContent() { return attachmentContent; }
    public void setAttachmentContent(String attachmentContent) { this.attachmentContent = attachmentContent; }
}

2. 结合Repository使用Pipeline

默认的ElasticsearchRepository.save()不会自动指定Pipeline,需要用ElasticsearchOperations来提交请求:

import org.springframework.data.elasticsearch.core.ElasticsearchOperations;
import org.springframework.data.elasticsearch.core.IndexedObjectInformation;
import org.springframework.stereotype.Service;
import java.io.File;
import java.io.IOException;
import java.nio.file.Files;

@Service
public class PdfDocService {

    private final ElasticsearchOperations elasticsearchOperations;
    private final PdfDocRepository pdfDocRepository;

    public PdfDocService(ElasticsearchOperations elasticsearchOperations, PdfDocRepository pdfDocRepository) {
        this.elasticsearchOperations = elasticsearchOperations;
        this.pdfDocRepository = pdfDocRepository;
    }

    public void indexPdf(File pdfFile) throws IOException {
        PdfDoc pdfDoc = new PdfDoc();
        pdfDoc.setFileName(pdfFile.getName());
        pdfDoc.setData(Files.readAllBytes(pdfFile.toPath()));

        // 指定Pipeline提交索引请求
        IndexedObjectInformation info = elasticsearchOperations.index(i -> i
                .index("pdf-docs")
                .pipeline("pdf-attachment-pipeline")
                .document(pdfDoc)
        );
        pdfDoc.setId(info.id()); // 可选:把ES生成的id赋值给实体
    }
}

对应的Repository接口:

import org.springframework.data.elasticsearch.repository.ElasticsearchRepository;

public interface PdfDocRepository extends ElasticsearchRepository<PdfDoc, String> {
}

四、实现PDF内容检索

直接针对Pipeline提取后的attachmentContent字段做全文检索即可:

1. 用Repository定义检索方法

import org.springframework.data.elasticsearch.repository.ElasticsearchRepository;
import java.util.List;

public interface PdfDocRepository extends ElasticsearchRepository<PdfDoc, String> {
    // 根据内容关键词检索
    List<PdfDoc> findByAttachmentContentContaining(String keyword);
}

2. 复杂查询用ElasticsearchOperations

import org.springframework.data.elasticsearch.core.SearchHits;
import org.springframework.data.elasticsearch.core.query.Criteria;
import org.springframework.data.elasticsearch.core.query.CriteriaQuery;
import java.util.List;
import java.util.stream.Collectors;

public List<PdfDoc> searchPdfContent(String keyword) {
    Criteria criteria = Criteria.where("attachmentContent").contains(keyword);
    CriteriaQuery query = new CriteriaQuery(criteria);
    SearchHits<PdfDoc> hits = elasticsearchOperations.search(query, PdfDoc.class);
    return hits.getSearchHits().stream()
            .map(hit -> hit.getContent())
            .collect(Collectors.toList());
}

注意事项

  • 确保Elasticsearch已安装ingest-attachment插件,安装命令:bin/elasticsearch-plugin install ingest-attachment
  • ES默认单文档大小限制100MB,若PDF超过此大小,需修改elasticsearch.yml中的http.max_content_length配置
  • 分词器要按需配置,中文推荐用IK分词器,需提前在ES中安装对应插件并在实体类字段上指定

内容的提问来源于stack exchange,提问作者TCSInd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 07:20:34