Spring Boot JPA批量同步10万条数据性能过慢求优化方案
数据同步效率问题分析与优化方案
一、耗时原因拆解
- 全量加载数据撑爆内存:直接调用
findAll()把十万级数据一次性加载到JVM内存,不仅占用大量堆空间,还会触发频繁的垃圾回收,拖慢整体执行速度。 - 批量插入逻辑错误:当前代码在每个
ResearchEntity循环后就调用saveAll,且researchInfoEntityList未清空,导致后续批次重复插入之前的数据,同时多次发起数据库连接请求,每次IO都产生额外开销。 - 未开启JPA批量插入支持:默认情况下JPA的
saveAll可能仍单条执行insert语句,未利用数据库的批量插入能力,每条插入都要走一次网络和事务流程。 - 大量临时对象增加GC负担:循环内每次都新建
ResearchInfoEntity实例,十万级数量下会产生大量临时对象,加剧GC压力。
二、优化后的批量插入代码
第一步:配置JPA批量插入参数
在application.yml(或application.properties)中添加以下配置,开启数据库批量插入支持:
spring: jpa: properties: hibernate: jdbc: batch_size: 500 # 批次大小,建议根据数据库性能调整为500-1000 order_inserts: true # 对insert语句排序,提升批量插入效率 order_updates: true
第二步:优化同步代码
// 分页加载PersonInfo数据,避免全量加载占用内存 int pageSize = 1000; Pageable pageable = PageRequest.of(0, pageSize); Map<Long, List<PersonInfoEntity>> personInfoMap = new HashMap<>(); // 分页查询并构建personId到PersonInfo的映射 while (true) { Page<PersonInfoEntity> personPage = personInfoRepository.findAll(pageable); if (personPage.isEmpty()) { break; } personPage.getContent().stream() .collect(Collectors.groupingBy(p -> p.getPerson().getPersonId())) .forEach((key, value) -> personInfoMap.merge(key, value, (v1, v2) -> { v1.addAll(v2); return v1; })); pageable = personPage.nextPageable(); } // 分页加载ResearchEntity,同步生成并批量插入ResearchInfo pageable = PageRequest.of(0, pageSize); int batchSize = 500; // 和hibernate.batch_size保持一致 List<ResearchInfoEntity> tempBatchList = new ArrayList<>(batchSize); while (true) { Page<ResearchEntity> researchPage = researchRepository.findAll(pageable); if (researchPage.isEmpty()) { break; } for (ResearchEntity researchEntity : researchPage.getContent()) { List<PersonInfoEntity> personList = personInfoMap.get(researchEntity.getPerson().getPersonId()); if (Objects.isNull(personList)) { continue; } for (PersonInfoEntity personInfo : personList) { ResearchInfoEntity researchInfo = new ResearchInfoEntity(); researchInfo.setRecovery(researchEntity); // 修正原代码的变量名错误 researchInfo.setMilestoneGroupId(personInfo.getMilestoneGroupId()); researchInfo.setMilestoneId(personInfo.getMilestoneId()); researchInfo.setMilestoneStepId(personInfo.getMilestoneStepId()); researchInfo.setMilestoneStepValue(personInfo.getMilestoneStepValue()); researchInfo.setCreateBy(personInfo.getCreateBy()); researchInfo.setCreateTime(personInfo.getCreateTime()); researchInfo.setUpdateBy(personInfo.getUpdateBy()); researchInfo.setUpdateTime(personInfo.getUpdateTime()); tempBatchList.add(researchInfo); // 达到批次大小就执行插入并清空临时列表 if (tempBatchList.size() >= batchSize) { researchInfoEntityRepository.saveAll(tempBatchList); tempBatchList.clear(); } } } // 处理剩余不足一个批次的数据 if (!tempBatchList.isEmpty()) { researchInfoEntityRepository.saveAll(tempBatchList); tempBatchList.clear(); } pageable = researchPage.nextPageable(); }
额外优化建议
- 用MapStruct简化实体映射:替换手动set属性,减少代码冗余,提升映射效率。
- 添加事务注解:在同步方法上添加
@Transactional,减少事务提交次数(注意事务范围不要过大,避免锁表)。 - 数据库连接池优化:调整数据库连接池参数,增加连接数提升并发插入能力。
内容的提问来源于stack exchange,提问作者user3684675
相关产品推荐
相关产品推荐

