Spring Batch中Reader到Processor再到Writer的数据传递机制咨询
Hey there! As someone who's worked with Spring Batch for a few years, let's break down exactly how data moves from your Reader to Processor—this is a core part of the framework's workflow, so it's a great question to nail down early.
Core Flow: Direct In-Memory Passing
First off, the short answer: Spring Batch doesn't persist the item from read() to some external store before passing it to process(). By default, the item moves directly through memory from the Reader to the Processor, with no intermediate persistent storage.
Let's walk through the step-by-step flow (using the default chunk-oriented processing model, which is what most Spring Batch jobs use):
- Step Execution Loop Starts
Your Step runs in a loop that continues untilreader.read()returnsnull(signaling no more data) or the job is stopped. - Single Item Read
Each iteration starts with callingreader.read(). If this returns a non-null item (like your String), the framework first runs any configured item-level interceptors (e.g.,AfterReadInterceptorif you set one up for validation or logging). - Direct Pass to Processor
Immediately after the read (and interceptor steps), the framework passes that single item directly intoprocessor.process(). No detours, no writing to a database or file—this is all in-memory. - Chunk Accumulation for Writer
The processed item gets added to an in-memory list. Once this list reaches your configuredchunk-size(e.g., 10 items), the entire list is passed towriter.write(). If the Processor returnsnullfor an item, that item is skipped and won't be added to the list.
When "Storage" Comes Into Play (The Exceptions)
The only time Spring Batch will persist an item between Reader and Processor is if you've configured skip or retry logic for failed items. For example:
- If
process()throws an exception you've marked as skippable, the framework will store the failed item (and its metadata) in the JobRepository's database tables (likeBATCH_STEP_EXECUTION_CONTEXT) so it can track which items were skipped. - Similarly, for retry, it might store state temporarily to handle retries without re-reading the entire dataset.
But for normal, successful item processing? No external storage—just straight in-memory handoff.
A Simplified "Framework Internals" Pseudocode
To make this even clearer, here's a rough approximation of what the framework is doing under the hood (super simplified, but captures the core flow):
List<Object> processedItems = new ArrayList<>(); int currentChunkCount = 0; Object nextItem; // Loop until no more items or chunk is full while ((nextItem = reader.read()) != null && currentChunkCount < chunkSize) { // Run any after-read interceptors try { Object processedItem = processor.process(nextItem); if (processedItem != null) { processedItems.add(processedItem); currentChunkCount++; } } catch (SkipException e) { // Handle skip logic, store item metadata if needed continue; } } // Once chunk is ready, write all processed items writer.write(processedItems);
Tips for Newbies to Solidify This
- Write a tiny demo job: Create a
ItemReader<String>that returns a list of strings, aItemProcessor<String, String>that converts them to uppercase, and aItemWriter<String>that prints the results. Debug through the code—you'll see each string go straight fromread()toprocess()without any stops. - Focus on chunk-oriented processing: This is the backbone of Spring Batch. Understanding how chunks balance performance (batch writes) and transactionality (each chunk is a single transaction) will answer most of your workflow questions.
内容的提问来源于stack exchange,提问作者Saurabh Deshpande

