Spring Batch配置分区后无法读取全部记录问题排查
问题分析与解决方案
核心问题
你的Spring Batch分区配置存在两个关键问题:
Reader未利用分区上下文参数过滤数据
分区器已经生成了每个分区的minValue和maxValue(id范围),但你的RepositoryItemReader没有从ExecutionContext中获取这些参数,也没有传递给Repository方法做范围过滤。这导致每个slave step的Reader都在读取全量符合条件的记录,既会重复处理数据,也会引发单分区时的读取异常。Reader未配置StepScope
没有给Reader添加@StepScope注解,导致Reader在容器初始化时就固定了参数,无法动态获取每个分区的上下文参数,这是分区场景下Reader必须的配置。
修复步骤
1. 修改Repository方法,添加id范围过滤参数
确保你的Repository接口新增带id范围的查询方法,与分区器的范围匹配:
// RecurringEarningRepository.java List<RecurringEarning> findRunnableRecurringEarnings(LocalDate date, Integer minId, Integer maxId);
对应的SQL或JPA查询需要添加id BETWEEN :minId AND :maxId的过滤条件。
2. 重构Reader配置,添加StepScope并获取分区参数
给Reader添加@StepScope注解,动态从stepExecutionContext中读取分区的min/max值,传递给Repository方法:
@Bean @StepScope // 必须添加,才能获取当前step的上下文参数 public RepositoryItemReader<RecurringEarning> reader( @Value("#{stepExecutionContext['minValue']}") Integer minValue, @Value("#{stepExecutionContext['maxValue']}") Integer maxValue) { RepositoryItemReader<RecurringEarning> reader = new RepositoryItemReader<>(); reader.setRepository(recurringEarningRepository); reader.setName("recurringEarningRepository"); reader.setMethodName("findRunnableRecurringEarnings"); // 传递日期+id范围参数 reader.setArguments(List.of(LocalDate.now(), minValue, maxValue)); reader.setPageSize(chunkSize); reader.setSort(Map.of("id", Sort.Direction.ASC)); return reader; }
3. 验证分区器的统计方法与查询方法逻辑一致
确保countValidRecords(LocalDate.now())的过滤条件和findRunnableRecurringEarnings完全相同,否则分区器计算的总记录数与实际可查询到的记录数不匹配,会导致分区范围划分错误。
额外注意事项
- 如果你的表中id不是连续的(存在删除记录导致id断层),基于id范围的分区会导致各分区处理的记录数差异较大,这种情况下可以考虑采用基于分页的分区器,或者先查询所有符合条件的id列表再拆分。
- 检查
chunkSize的设置,确保分页读取能覆盖分区内的所有数据。
内容的提问来源于stack exchange,提问作者Sithira
相关产品推荐
相关产品推荐

