You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Spring Batch Job中使用CSVPrinter实现CSV文件追加写入

解决方案

问题根源

  1. 每次调用write方法处理批次数据时,resource.getOutputStream()默认是覆盖写入模式,新批次数据会直接覆盖旧文件内容,导致最终仅保留最后一批数据
  2. 如果直接改成追加模式,还会出现每个批次都重复写入CSV表头的问题,不符合CSV格式规范

首选方案(改动最小,兼容所有存储类型)

如果数据量未超出JVM内存上限,推荐采用缓存全量数据、Step执行完成后一次性写入的方案,仅需要少量代码调整即可解决问题,原有printCSV逻辑完全不需要修改:

@Component
@EnableContextResourceLoader
@StepScope // 加StepScope保证每个Step对应单独的Writer实例,避免多任务数据混淆
class LocationWriter implements ItemWriter<Location> {
  @Autowired
  ResourceLoader resourceLoader

  @Resource
  FileProperties fileProperties

  // 缓存全量位置数据
  private List<Location> allLocations = []

  @BeforeStep
  void beforeStep(StepExecution stepExecution) {
    allLocations.clear()
  }

  @Override
  void write(List<? extends Location> items) throws Exception {
    // 批次数据暂存到内存,不直接写文件
    allLocations.addAll(items)
  }

  @AfterStep
  ExitStatus afterStep(StepExecution stepExecution) {
    if (allLocations.isEmpty()) {
      return stepExecution.getExitStatus()
    }
    // 整个Step执行完成后一次性写入全量数据
    WritableResource resource = resourceLoader.getResource("s3://${fileProperties.bucket}/${fileProperties.outputLocation}")
    try (OutputStreamWriter writer = new OutputStreamWriter(resource.getOutputStream(), StandardCharsets.UTF_8)) {
      printCSV(writer, allLocations)
    } catch (Exception e) {
      stepExecution.addFailureException(e)
      return ExitStatus.FAILED
    }
    return stepExecution.getExitStatus()
  }

  static void printCSV(OutputStreamWriter writer, List<Location> summaries) {
    new CSVPrinter(
      writer,
      CSVFormat.DEFAULT
        .withDelimiter('|' as char)
        .withHeader("id", "address", "city")
    ).withCloseable { CSVPrinter csvPrinter -> 
      summaries.each { Location location ->
        csvPrinter.printRecord(
          location.id,
          location.address,
          location.city
        )   
      }
    }
  }
}

方案优势

  • 仅新增3个方法/变量,原有打印逻辑完全复用,改动量极小
  • 天然避免覆盖写入、重复表头问题
  • 兼容S3等不支持追加写入的存储服务

备选方案(大数据量适用,仅支持可追加存储)

如果数据量过大无法全量缓存到内存,且你使用的存储支持追加写入(如本地文件系统),可以采用追加写入+控制表头仅输出一次的方案:

@Component
@EnableContextResourceLoader
@StepScope
class LocationWriter implements ItemWriter<Location> {
  @Autowired
  ResourceLoader resourceLoader

  @Resource
  FileProperties fileProperties

  // 标记是否是第一次写入,控制表头仅输出一次
  private volatile boolean firstWrite = true

  @Override
  void write(List<? extends Location> items) throws Exception {
    WritableResource resource = resourceLoader.getResource("file:${fileProperties.outputPath}")
    // 以追加模式打开输出流
    OutputStream os = new FileOutputStream(resource.getFile(), true)
    try (OutputStreamWriter writer = new OutputStreamWriter(os, StandardCharsets.UTF_8)) {
      printCSV(writer, items, firstWrite)
      firstWrite = false
    }
  }

  static void printCSV(OutputStreamWriter writer, List<Location> summaries, boolean writeHeader) {
    CSVFormat format = CSVFormat.DEFAULT.withDelimiter('|' as char)
    // 仅第一次写入时输出表头
    if (writeHeader) {
      format = format.withHeader("id", "address", "city")
    }
    new CSVPrinter(writer, format).withCloseable { CSVPrinter csvPrinter -> 
      summaries.each { Location location ->
        csvPrinter.printRecord(
          location.id,
          location.address,
          location.city
        )   
      }
    }
  }
}

注意事项

S3对象存储原生不支持追加写入,该方案不适用于S3存储,大数据量写入S3需要采用分块写多个临时文件、最后合并为单一文件的实现方式。

内容的提问来源于stack exchange,提问作者p4rt2020

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 19:27:03