Spring Batch中关联处理城市与历史文件并统一写入JSON的疑问
解决方案:Spring Batch 城市与历史数据关联处理建议
针对你遇到的城市与历史数据关联需求,推荐采用「先加载历史数据到缓存,再处理城市并关联」的两步结构,而非合并为单个Step。这种方式既符合Spring Batch职责单一的设计原则,也便于后续维护和扩展,同时能解决数据关联的核心需求。
核心思路
- 调整Step执行顺序:先执行历史数据加载Step,将所有历史记录按邮政编码(匹配键)存入缓存
- 处理城市数据时,在Processor中从缓存匹配对应历史记录,关联到City对象
- 使用单个ItemWriter写入关联后的完整City数据,避免JSON结构混乱
具体代码实现
1. 创建历史数据缓存Bean
用于在Step间共享历史数据:
@Component public class HistoriqueCache { private final Map<String, List<Historique>> historiqueMap = new ConcurrentHashMap<>(); public void put(String postalCode, Historique historique) { historiqueMap.computeIfAbsent(postalCode, k -> new ArrayList<>()).add(historique); } public List<Historique> get(String postalCode) { return historiqueMap.getOrDefault(postalCode, Collections.emptyList()); } }
2. 修改历史数据Step(仅加载缓存,无需写入文件)
@Bean public Step stepLoadHistorique(JobRepository jobRepository, PlatformTransactionManager txManager, HistoriqueCache historiqueCache) throws Exception { return new StepBuilder("stepLoadHistorique", jobRepository) .<Historique, Historique>chunk(25, txManager) .reader(readerHistorique()) .processor(historique -> { // 按实际匹配规则取邮政编码(示例用com_ap作为匹配键) String postalCode = historique.getCom_ap(); historiqueCache.put(postalCode, historique); return historique; }) .writer(items -> { /* 无需写入文件,仅加载缓存 */ }) .build(); }
3. 修改城市Step的Processor,关联历史数据
确保City实体类添加List<Historique> historiques属性及对应的setter方法:
@Bean public ItemProcessor<City, City> processorCity(HistoriqueCache historiqueCache) { return city -> { // 用City的com字段作为匹配键,关联对应历史记录 List<Historique> matchedHistoriques = historiqueCache.get(city.getCom()); city.setHistoriques(matchedHistoriques); // 保留原有的城市数据处理逻辑 return city; }; }
4. 调整Job的Step执行顺序
@Bean public Job runjob(JobRepository jobRepository, PlatformTransactionManager txManager, Step stepLoadHistorique, Step stepCity) throws Exception{ return new JobBuilder("runJob", jobRepository) .start(stepCountry(jobRepository,txManager)) .next(stepRegion(jobRepository,txManager)) .next(stepDepartment(jobRepository,txManager)) .next(stepLoadHistorique) // 先加载历史缓存 .next(stepCity) // 再处理城市并关联 .build(); }
5. 统一使用单个ItemWriter写入关联数据
删除原有的writerHistorique,修改writerCity以支持写入包含历史记录的City对象:
@Bean public ItemWriter<City> writerCity() { JsonFileItemWriter<City> writerCity = new JsonFileItemWriter<>( new FileSystemResource("src/main/java/output/Insee.json"), new JsonObjectMarshaller<City>() { @Override public String marshal(City object) { try { ObjectMapper mapper = new ObjectMapper(); mapper.configure(SerializationFeature.INDENT_OUTPUT, true); return mapper.writeValueAsString(object); } catch (JsonProcessingException e) { e.printStackTrace(); return null; } } } ); writerCity.setAppendAllowed(false); // 重新生成完整的JSON文件,避免结构混乱 return writerCity; }
为什么不推荐合并为单个Step?
- 合并Step需要同时处理两个文件的读取与匹配,Chunk处理模式下难以保证数据匹配的完整性,逻辑复杂度大幅提升
- 分步处理职责清晰:一个Step负责加载数据,一个负责业务处理与关联,便于调试和后续功能扩展
- 如果历史文件体积过大,缓存方案可轻松替换为临时数据库存储,分步结构无需大幅修改
内容的提问来源于stack exchange,提问作者mat
相关产品推荐
相关产品推荐

