You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spring Boot JPA批量同步10万条数据性能过慢求优化方案

数据同步效率问题分析与优化方案

一、耗时原因拆解

  1. 全量加载数据撑爆内存:直接调用findAll()把十万级数据一次性加载到JVM内存,不仅占用大量堆空间,还会触发频繁的垃圾回收,拖慢整体执行速度。
  2. 批量插入逻辑错误:当前代码在每个ResearchEntity循环后就调用saveAll,且researchInfoEntityList未清空,导致后续批次重复插入之前的数据,同时多次发起数据库连接请求,每次IO都产生额外开销。
  3. 未开启JPA批量插入支持:默认情况下JPA的saveAll可能仍单条执行insert语句,未利用数据库的批量插入能力,每条插入都要走一次网络和事务流程。
  4. 大量临时对象增加GC负担:循环内每次都新建ResearchInfoEntity实例,十万级数量下会产生大量临时对象,加剧GC压力。

二、优化后的批量插入代码

第一步:配置JPA批量插入参数

在application.yml(或application.properties)中添加以下配置,开启数据库批量插入支持:

spring:
  jpa:
    properties:
      hibernate:
        jdbc:
          batch_size: 500  # 批次大小,建议根据数据库性能调整为500-1000
        order_inserts: true  # 对insert语句排序,提升批量插入效率
        order_updates: true

第二步:优化同步代码

// 分页加载PersonInfo数据,避免全量加载占用内存
int pageSize = 1000;
Pageable pageable = PageRequest.of(0, pageSize);
Map<Long, List<PersonInfoEntity>> personInfoMap = new HashMap<>();

// 分页查询并构建personId到PersonInfo的映射
while (true) {
    Page<PersonInfoEntity> personPage = personInfoRepository.findAll(pageable);
    if (personPage.isEmpty()) {
        break;
    }
    personPage.getContent().stream()
            .collect(Collectors.groupingBy(p -> p.getPerson().getPersonId()))
            .forEach((key, value) -> personInfoMap.merge(key, value, (v1, v2) -> {
                v1.addAll(v2);
                return v1;
            }));
    pageable = personPage.nextPageable();
}

// 分页加载ResearchEntity,同步生成并批量插入ResearchInfo
pageable = PageRequest.of(0, pageSize);
int batchSize = 500; // 和hibernate.batch_size保持一致
List<ResearchInfoEntity> tempBatchList = new ArrayList<>(batchSize);

while (true) {
    Page<ResearchEntity> researchPage = researchRepository.findAll(pageable);
    if (researchPage.isEmpty()) {
        break;
    }

    for (ResearchEntity researchEntity : researchPage.getContent()) {
        List<PersonInfoEntity> personList = personInfoMap.get(researchEntity.getPerson().getPersonId());
        if (Objects.isNull(personList)) {
            continue;
        }

        for (PersonInfoEntity personInfo : personList) {
            ResearchInfoEntity researchInfo = new ResearchInfoEntity();
            researchInfo.setRecovery(researchEntity); // 修正原代码的变量名错误
            researchInfo.setMilestoneGroupId(personInfo.getMilestoneGroupId());
            researchInfo.setMilestoneId(personInfo.getMilestoneId());
            researchInfo.setMilestoneStepId(personInfo.getMilestoneStepId());
            researchInfo.setMilestoneStepValue(personInfo.getMilestoneStepValue());
            researchInfo.setCreateBy(personInfo.getCreateBy());
            researchInfo.setCreateTime(personInfo.getCreateTime());
            researchInfo.setUpdateBy(personInfo.getUpdateBy());
            researchInfo.setUpdateTime(personInfo.getUpdateTime());

            tempBatchList.add(researchInfo);

            // 达到批次大小就执行插入并清空临时列表
            if (tempBatchList.size() >= batchSize) {
                researchInfoEntityRepository.saveAll(tempBatchList);
                tempBatchList.clear();
            }
        }
    }

    // 处理剩余不足一个批次的数据
    if (!tempBatchList.isEmpty()) {
        researchInfoEntityRepository.saveAll(tempBatchList);
        tempBatchList.clear();
    }

    pageable = researchPage.nextPageable();
}

额外优化建议

  • 用MapStruct简化实体映射:替换手动set属性,减少代码冗余,提升映射效率。
  • 添加事务注解:在同步方法上添加@Transactional,减少事务提交次数(注意事务范围不要过大,避免锁表)。
  • 数据库连接池优化:调整数据库连接池参数,增加连接数提升并发插入能力。

内容的提问来源于stack exchange,提问作者user3684675

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 01:10:27