Spring Batch CompositeItemProcessor委托间数据共享方案咨询
Let's break down your problem and walk through the right ways to share data between your ItemProcessor delegates in a CompositeItemProcessor, plus address your question about using a util:map.
Why Your Previous Attempts Didn't Work
First, let's clarify why the approaches you tried fell short:
- Step Execution Context: This is shared across the entire step (and persisted across commits). For a single chunk (especially with
commit-interval=1), it might seem like it could work, but it's not designed for per-item intermediate data—if you have a multi-threaded step, multiple items will overwrite values in the context, leading to race conditions. Even in single-threaded mode, it's not the right tool for passing temporary per-item data between processors. @AfterProcess: This annotation's method runs after the entireItemProcessor(all delegates) has finished processing the item, not after your individualItemProcessor1completes. That's why it was executing too late for your needs.
Recommended Ways to Share Data Between Delegates
1. Modify the Item (or Use a Wrapper Class)
If you can adjust your item structure, this is the cleanest and most thread-safe approach. Create a wrapper class that holds your original item plus a map of shared intermediate values. All your processors will work with this wrapper, passing along the shared data as they process the item.
Example wrapper class:
public class ItemWrapper<T> { private T originalItem; private Map<String, Object> sharedMetadata = new HashMap<>(); // Getters and setters for both fields }
Then update your processors to handle the wrapper:
// ItemProcessor1 public class ItemProcessor1 implements ItemProcessor<ItemWrapper<InputItem>, ItemWrapper<IntermediateItem1>> { @Override public ItemWrapper<IntermediateItem1> process(ItemWrapper<InputItem> wrapper) throws Exception { InputItem input = wrapper.getOriginalItem(); // Process input to get IntermediateItem1 IntermediateItem1 result = ...; // Store your value in the shared metadata wrapper.getSharedMetadata().put("valueFromProcessor1", yourValueHere); // Create a new wrapper for the next processor, carrying over the shared metadata ItemWrapper<IntermediateItem1> nextWrapper = new ItemWrapper<>(); nextWrapper.setOriginalItem(result); nextWrapper.setSharedMetadata(wrapper.getSharedMetadata()); return nextWrapper; } }
In ItemProcessor4, you can then retrieve the values directly from the wrapper's shared metadata:
public class ItemProcessor4 implements ItemProcessor<ItemWrapper<IntermediateItem3>, OutputItem> { @Override public OutputItem process(ItemWrapper<IntermediateItem3> wrapper) throws Exception { Object valueFrom1 = wrapper.getSharedMetadata().get("valueFromProcessor1"); Object valueFrom3 = wrapper.getSharedMetadata().get("valueFromProcessor3"); // Use these values to process and return your OutputItem return ...; } }
This approach is inherently thread-safe because each item gets its own wrapper instance—no cross-item or cross-thread interference, regardless of commit-interval or step threading.
2. Use Thread-Local Storage
If you can't modify the item structure, thread-local storage is the next best option. Since each item is processed in a single thread (especially with commit-interval=1), you can use a ThreadLocal to store per-thread/per-item intermediate data.
First, create a helper class to manage the thread-local data:
public class ProcessorSharedData { private static final ThreadLocal<Map<String, Object>> threadLocalData = new ThreadLocal<>(); // Initialize the storage for a new item processing cycle public static void init() { threadLocalData.set(new HashMap<>()); } // Store a value public static void put(String key, Object value) { threadLocalData.get().put(key, value); } // Retrieve a value public static Object get(String key) { return threadLocalData.get().get(key); } // Clean up to avoid memory leaks public static void clear() { threadLocalData.remove(); } }
Then update your processors:
// ItemProcessor1 public class ItemProcessor1 implements ItemProcessor<InputItem, IntermediateItem1> { @Override public IntermediateItem1 process(InputItem item) throws Exception { ProcessorSharedData.init(); // Initialize for this item // Process item to get result IntermediateItem1 result = ...; // Store your value ProcessorSharedData.put("valueFromProcessor1", yourValueHere); return result; } } // ItemProcessor3 (similar to Processor1) public class ItemProcessor3 implements ItemProcessor<IntermediateItem2, IntermediateItem3> { @Override public IntermediateItem3 process(IntermediateItem2 item) throws Exception { // Process item to get result IntermediateItem3 result = ...; // Store your value ProcessorSharedData.put("valueFromProcessor3", yourValueHere); return result; } } // ItemProcessor4 public class ItemProcessor4 implements ItemProcessor<IntermediateItem3, OutputItem> { @Override public OutputItem process(IntermediateItem3 item) throws Exception { // Retrieve values Object valueFrom1 = ProcessorSharedData.get("valueFromProcessor1"); Object valueFrom3 = ProcessorSharedData.get("valueFromProcessor3"); // Process to get output OutputItem result = ...; // Clean up the thread-local storage to prevent leaks ProcessorSharedData.clear(); return result; } }
Important notes for this approach:
- Always call
clear()after processing the item to avoid memory leaks, especially if your thread pool reuses threads. - This works perfectly with
commit-interval=1(each item gets its own thread-local context). If you ever increasecommit-interval, you'll need to adjust the cleanup to happen after the entire chunk is processed (e.g., using aChunkListener).
Your Question: Using util:map with commit-interval=1?
If you're thinking of using a singleton HashMap (via util:map), this is not safe unless your step is strictly single-threaded. Even with commit-interval=1, if you ever switch to a multi-threaded TaskExecutor, multiple threads will overwrite values in the map, leading to corrupted data.
If you must use a map, wrap it in a ThreadLocal (like the helper class in option 2) to ensure each thread has its own isolated map. This way, you avoid race conditions while still using a map structure.
内容的提问来源于stack exchange,提问作者Robert Mason

