如何在Spring Batch Job中使用CSVPrinter实现CSV文件追加写入
解决方案
问题根源
- 每次调用
write方法处理批次数据时,resource.getOutputStream()默认是覆盖写入模式,新批次数据会直接覆盖旧文件内容,导致最终仅保留最后一批数据 - 如果直接改成追加模式,还会出现每个批次都重复写入CSV表头的问题,不符合CSV格式规范
首选方案(改动最小,兼容所有存储类型)
如果数据量未超出JVM内存上限,推荐采用缓存全量数据、Step执行完成后一次性写入的方案,仅需要少量代码调整即可解决问题,原有printCSV逻辑完全不需要修改:
@Component @EnableContextResourceLoader @StepScope // 加StepScope保证每个Step对应单独的Writer实例,避免多任务数据混淆 class LocationWriter implements ItemWriter<Location> { @Autowired ResourceLoader resourceLoader @Resource FileProperties fileProperties // 缓存全量位置数据 private List<Location> allLocations = [] @BeforeStep void beforeStep(StepExecution stepExecution) { allLocations.clear() } @Override void write(List<? extends Location> items) throws Exception { // 批次数据暂存到内存,不直接写文件 allLocations.addAll(items) } @AfterStep ExitStatus afterStep(StepExecution stepExecution) { if (allLocations.isEmpty()) { return stepExecution.getExitStatus() } // 整个Step执行完成后一次性写入全量数据 WritableResource resource = resourceLoader.getResource("s3://${fileProperties.bucket}/${fileProperties.outputLocation}") try (OutputStreamWriter writer = new OutputStreamWriter(resource.getOutputStream(), StandardCharsets.UTF_8)) { printCSV(writer, allLocations) } catch (Exception e) { stepExecution.addFailureException(e) return ExitStatus.FAILED } return stepExecution.getExitStatus() } static void printCSV(OutputStreamWriter writer, List<Location> summaries) { new CSVPrinter( writer, CSVFormat.DEFAULT .withDelimiter('|' as char) .withHeader("id", "address", "city") ).withCloseable { CSVPrinter csvPrinter -> summaries.each { Location location -> csvPrinter.printRecord( location.id, location.address, location.city ) } } } }
方案优势
- 仅新增3个方法/变量,原有打印逻辑完全复用,改动量极小
- 天然避免覆盖写入、重复表头问题
- 兼容S3等不支持追加写入的存储服务
备选方案(大数据量适用,仅支持可追加存储)
如果数据量过大无法全量缓存到内存,且你使用的存储支持追加写入(如本地文件系统),可以采用追加写入+控制表头仅输出一次的方案:
@Component @EnableContextResourceLoader @StepScope class LocationWriter implements ItemWriter<Location> { @Autowired ResourceLoader resourceLoader @Resource FileProperties fileProperties // 标记是否是第一次写入,控制表头仅输出一次 private volatile boolean firstWrite = true @Override void write(List<? extends Location> items) throws Exception { WritableResource resource = resourceLoader.getResource("file:${fileProperties.outputPath}") // 以追加模式打开输出流 OutputStream os = new FileOutputStream(resource.getFile(), true) try (OutputStreamWriter writer = new OutputStreamWriter(os, StandardCharsets.UTF_8)) { printCSV(writer, items, firstWrite) firstWrite = false } } static void printCSV(OutputStreamWriter writer, List<Location> summaries, boolean writeHeader) { CSVFormat format = CSVFormat.DEFAULT.withDelimiter('|' as char) // 仅第一次写入时输出表头 if (writeHeader) { format = format.withHeader("id", "address", "city") } new CSVPrinter(writer, format).withCloseable { CSVPrinter csvPrinter -> summaries.each { Location location -> csvPrinter.printRecord( location.id, location.address, location.city ) } } } }
注意事项
S3对象存储原生不支持追加写入,该方案不适用于S3存储,大数据量写入S3需要采用分块写多个临时文件、最后合并为单一文件的实现方式。
内容的提问来源于stack exchange,提问作者p4rt2020
相关产品推荐
相关产品推荐

