You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spring Batch中MultiResourceItemReader单文件读取后TaskExecutor多线程处理问题

Hey, I see what's going on here. Your current setup is making Spring Batch process chunks from multiple files in parallel, which isn't what you want. Let's break down why this happens and how to fix it.

The Problem with Your Current Configuration

The issue with your existing Step1 is that you've added a taskExecutor() directly to the step. When you do this, Spring Batch parallelizes chunk processing across the entire step. Since MultiResourceItemReader fetches chunks from all your CSV files at once, those chunks get distributed to different threads—resulting in multiple files being processed simultaneously.

Your goal is to process files one after another (serially at the file level), but process the records within each file in parallel. To achieve this, we need to split the work into two layers:

  1. A master step that iterates through your files one by one.
  2. A slave step that handles parallel processing of records for each individual file.

The cleanest way to do this in Spring Batch is using partitioning. Here's how to set it up:


1. Create a Partitioner for Your Files

First, we need a partitioner that creates a separate "partition" for each CSV file. Each partition will hold the path to one file.

@Bean
public Partitioner filePartitioner() {
    return gridSize -> {
        Map<String, ExecutionContext> partitions = new HashMap<>();
        try {
            Resource[] resources = resolver.getResources("file:/temp/*.csv");
            int partitionIndex = 0;
            for (Resource resource : resources) {
                ExecutionContext context = new ExecutionContext();
                context.putString("filePath", resource.getFile().getAbsolutePath());
                partitions.put("partition-" + partitionIndex, context);
                partitionIndex++;
            }
        } catch (IOException e) {
            throw new RuntimeException("Failed to load CSV files", e);
        }
        return partitions;
    };
}

2. Build a Slave Step for Parallel Record Processing

This step will handle processing records from a single file, using your task executor to parallelize chunk processing within that file. We'll use @StepScope to inject the file path from the partition's context.

@Bean
public Step slaveStep() {
    return stepBuilderFactory.get("slaveStep")
            .<POJO, POJO>chunk(5)
            .reader(singleFileItemReader(null))
            .writer(writer())
            .taskExecutor(taskExecutor())
            .throttleLimit(throttleLimit)
            .build();
}

@Bean
@StepScope
public FlatFileItemReader<POJO> singleFileItemReader(
        @Value("#{stepExecutionContext['filePath']}") String filePath) {
    FlatFileItemReader<POJO> reader = new FlatFileItemReader<>();
    reader.setResource(new FileSystemResource(filePath));
    // Add your existing LineMapper configuration here (same as your old itemReader())
    reader.setLineMapper(lineMapper());
    return reader;
}

3. Create a Master Step to Run Partitions Serially

The master step will use the partitioner to iterate through each file one by one. We don't add a task executor here—this ensures partitions (files) are processed in sequence.

@Bean
public Step masterStep() {
    return stepBuilderFactory.get("masterStep")
            .partitioner(slaveStep().getName(), filePartitioner())
            .step(slaveStep())
            .build();
}

4. Update Your Job Configuration

Finally, wire up your job to use the master step instead of your original Step1:

@Bean
public Job fileProcessingJob() {
    return jobBuilderFactory.get("fileProcessingJob")
            .start(masterStep())
            .build();
}

Why This Works

  • The master step runs partitions (files) one after another (serially) because we don't attach a task executor to it.
  • Each slave step processes a single file, using your task executor to parallelize chunk processing only within that file.
  • This exactly matches your desired behavior: process File 1's records in parallel → finish File 1 → process File 2's records in parallel → finish File 2 → process File 3.

Alternative Approach (If Partitioning Feels Overkill)

If you prefer a simpler setup, you could use two steps:

  1. A single-threaded step that reads all file paths and stores them in the job execution context.
  2. A loop step that processes each file path one by one, using a parallel step for each file. But partitioning is the more idiomatic Spring Batch way to handle this scenario.

内容的提问来源于stack exchange,提问作者drp111

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 17:17:47