You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spring Batch中关联处理城市与历史文件并统一写入JSON的疑问

解决方案:Spring Batch 城市与历史数据关联处理建议

针对你遇到的城市与历史数据关联需求,推荐采用「先加载历史数据到缓存,再处理城市并关联」的两步结构,而非合并为单个Step。这种方式既符合Spring Batch职责单一的设计原则,也便于后续维护和扩展,同时能解决数据关联的核心需求。

核心思路

  1. 调整Step执行顺序:先执行历史数据加载Step,将所有历史记录按邮政编码(匹配键)存入缓存
  2. 处理城市数据时,在Processor中从缓存匹配对应历史记录,关联到City对象
  3. 使用单个ItemWriter写入关联后的完整City数据,避免JSON结构混乱

具体代码实现

1. 创建历史数据缓存Bean

用于在Step间共享历史数据:

@Component
public class HistoriqueCache {
    private final Map<String, List<Historique>> historiqueMap = new ConcurrentHashMap<>();

    public void put(String postalCode, Historique historique) {
        historiqueMap.computeIfAbsent(postalCode, k -> new ArrayList<>()).add(historique);
    }

    public List<Historique> get(String postalCode) {
        return historiqueMap.getOrDefault(postalCode, Collections.emptyList());
    }
}

2. 修改历史数据Step(仅加载缓存,无需写入文件)

@Bean
public Step stepLoadHistorique(JobRepository jobRepository, PlatformTransactionManager txManager, HistoriqueCache historiqueCache) throws Exception {
    return new StepBuilder("stepLoadHistorique", jobRepository)
            .<Historique, Historique>chunk(25, txManager)
            .reader(readerHistorique())
            .processor(historique -> {
                // 按实际匹配规则取邮政编码(示例用com_ap作为匹配键)
                String postalCode = historique.getCom_ap();
                historiqueCache.put(postalCode, historique);
                return historique;
            })
            .writer(items -> { /* 无需写入文件,仅加载缓存 */ })
            .build();
}

3. 修改城市Step的Processor,关联历史数据

确保City实体类添加List<Historique> historiques属性及对应的setter方法:

@Bean
public ItemProcessor<City, City> processorCity(HistoriqueCache historiqueCache) {
    return city -> {
        // 用City的com字段作为匹配键,关联对应历史记录
        List<Historique> matchedHistoriques = historiqueCache.get(city.getCom());
        city.setHistoriques(matchedHistoriques);
        // 保留原有的城市数据处理逻辑
        return city;
    };
}

4. 调整Job的Step执行顺序

@Bean
public Job runjob(JobRepository jobRepository, PlatformTransactionManager txManager, Step stepLoadHistorique, Step stepCity) throws Exception{
    return new JobBuilder("runJob", jobRepository)
            .start(stepCountry(jobRepository,txManager))
            .next(stepRegion(jobRepository,txManager))
            .next(stepDepartment(jobRepository,txManager))
            .next(stepLoadHistorique) // 先加载历史缓存
            .next(stepCity) // 再处理城市并关联
            .build(); 
}

5. 统一使用单个ItemWriter写入关联数据

删除原有的writerHistorique,修改writerCity以支持写入包含历史记录的City对象:

@Bean
public ItemWriter<City> writerCity() {
    JsonFileItemWriter<City> writerCity = new JsonFileItemWriter<>(
        new FileSystemResource("src/main/java/output/Insee.json"),
        new JsonObjectMarshaller<City>() {
            @Override
            public String marshal(City object) {
                try {
                    ObjectMapper mapper = new ObjectMapper();
                    mapper.configure(SerializationFeature.INDENT_OUTPUT, true);
                    return mapper.writeValueAsString(object);
                } catch (JsonProcessingException e) {
                    e.printStackTrace();
                    return null;
                }
            }
        }
    );
    writerCity.setAppendAllowed(false); // 重新生成完整的JSON文件,避免结构混乱
    return writerCity;
}

为什么不推荐合并为单个Step?

  • 合并Step需要同时处理两个文件的读取与匹配,Chunk处理模式下难以保证数据匹配的完整性,逻辑复杂度大幅提升
  • 分步处理职责清晰:一个Step负责加载数据,一个负责业务处理与关联,便于调试和后续功能扩展
  • 如果历史文件体积过大,缓存方案可轻松替换为临时数据库存储,分步结构无需大幅修改

内容的提问来源于stack exchange,提问作者mat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 06:25:08