You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spring Batch步骤间大量ID列表(21000条)传递及业务实现求助

Spring Batch Step Communication: Accessing StepExecutionContext in Chunk Steps

Problem Scenario

  • Query 21,000 IDs from a database table to populate a List<Id>
  • In an existing chunk-based Step, remove elements present in this list within its ItemWriter
  • If the remaining list is not empty, execute an update on the table using those IDs

You've implemented a Tasklet Step to fetch the IDs and store them in the StepExecutionContext:

public RepeatStatus execute(StepContribution contribution, ChunkContext chunkContext) throws Exception {
    // 从数据库查询数据并创建Acct列表
    List acct = jdbcTemplate.query("SELECT id FROM Account",
        new BeanPropertyRowMapper<>(Acct.class));

    // 将列表存储到StepExecutionContext以传递给后续步骤
    chunkContext.getStepContext().getStepExecution().getExecutionContext().put("acct", acct);

    return RepeatStatus.FINISHED;
}

Core Question

How to retrieve this list in the chunk-based step, and what are reliable step communication solutions?

Solutions

1. Accessing the Stored List in Chunk Components

There are two straightforward ways to get the list from the StepExecutionContext in your ItemWriter:

Option A: Implement StepExecutionListener

Have your ItemWriter implement StepExecutionListener to gain access to the StepExecution object:

public class CustomItemWriter implements ItemWriter<YourEntity>, StepExecutionListener {
    private List<Acct> acctList;

    @Override
    public void beforeStep(StepExecution stepExecution) {
        // 从StepExecutionContext中取出存储的列表
        this.acctList = (List<Acct>) stepExecution.getExecutionContext().get("acct");
    }

    @Override
    public void write(List<? extends YourEntity> items) throws Exception {
        // 移除列表中存在的元素
        items.removeIf(item -> acctList.stream().anyMatch(acct -> acct.getId().equals(item.getId())));
        
        // 处理剩余元素的写入逻辑
        // ...
        
        // 若剩余元素不为空,执行更新操作
        if (!items.isEmpty()) {
            // 此处编写你的数据库更新逻辑
            // jdbcTemplate.update("UPDATE your_table SET ... WHERE id = ?", itemIds);
        }
    }

    @Override
    public ExitStatus afterStep(StepExecution stepExecution) {
        return null;
    }
}

Don't forget to register the listener in your step configuration:

@Bean
public Step chunkStep(ItemReader<YourEntity> reader, CustomItemWriter writer) {
    return stepBuilderFactory.get("chunkStep")
            .<YourEntity, YourEntity>chunk(100)
            .reader(reader)
            .writer(writer)
            .listener(writer) // 注册监听器
            .build();
}

Option B: Use @BeforeStep Annotation

Skip explicit listener implementation and use Spring Batch's @BeforeStep annotation to inject the StepExecution directly:

public class CustomItemWriter implements ItemWriter<YourEntity> {
    private List<Acct> acctList;

    @BeforeStep
    public void retrieveAcctList(StepExecution stepExecution) {
        this.acctList = (List<Acct>) stepExecution.getExecutionContext().get("acct");
    }

    @Override
    public void write(List<? extends YourEntity> items) throws Exception {
        // 与Option A相同的过滤和更新逻辑
        items.removeIf(item -> acctList.stream().anyMatch(acct -> acct.getId().equals(item.getId())));
        
        if (!items.isEmpty()) {
            // 执行更新操作
        }
    }
}

Spring Batch will automatically handle injecting the StepExecution without extra configuration.

2. Cross-Step Communication Best Practices

  • Use JobExecutionContext for multi-step sharing: If you need to pass data between more than two steps, store it in the JobExecutionContext instead. Modify your Tasklet code like this:

    chunkContext.getStepContext().getStepExecution().getJobExecution().getExecutionContext().put("acct", acct);
    

    Retrieve it the same way in any step component.

  • Ensure serializability: Objects stored in execution contexts must implement Serializable (your Acct class should follow this), as Spring Batch may persist the context to the database.

  • Avoid oversized datasets: While 21k entries are manageable, extremely large lists can cause memory issues. For million-scale records, consider using temporary database tables instead of in-memory lists.

内容的提问来源于stack exchange,提问作者ashwin bodkhe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 04:22:14