Spring Batch作业步骤无法完整读取大CSV文件问题排查求助
Hey there, let's dig into why your Spring Batch job wasn't processing all 13 million records from that 2.5GB CSV, and how you can fix this kind of issue in the future:
Possible Root Causes
Looking at your code and problem description, here are the most likely culprits:
- Thread-safety issues with your custom reader:
FlatFileItemReaderis not thread-safe by default. When you used aTaskExecutorfor parallel processing, multiple threads were accessing the same reader instance, leading to skipped records, duplicate reads, or incomplete file traversal. This is almost certainly the main reason only ~400k records were processed instead of 13 million. - Unintended breakage of default reader logic: By extending
FlatFileItemReaderand overridingafterPropertiesSet(), you might have accidentally overridden or skipped some of Spring Batch's built-in logic for handling large files, like proper stream management or error recovery. - Potential transaction/timeout issues: If your chunk size was misconfigured or your transaction manager had a low timeout setting, some chunks might have rolled back without you noticing. Though this is less likely given the huge discrepancy in record counts.
Fixes & Best Practices
1. Ensure Thread Safety for Parallel Processing
If you need to use multi-threading, wrap your reader in a SynchronizedItemStreamReader to enforce thread-safe access:
@Bean public ItemReader<ExperianPortal> threadSafeReader(MyReader myReader) { SynchronizedItemStreamReader<ExperianPortal> synchronizedReader = new SynchronizedItemStreamReader<>(); synchronizedReader.setDelegate(myReader); return synchronizedReader; }
Then use this wrapped reader in your step definition instead of the raw MyReader. For even better scalability with large files, consider using partitioning (PartitionStep) to split the CSV into smaller chunks, each processed by an independent reader instance.
2. Prioritize Spring Batch's Standard Components
As you discovered, switching to default Spring Batch components resolved the issue. That's because these components are battle-tested for edge cases like large files and parallel processing. If you need custom logic, prefer composition over inheritance:
- Customize the
LineMapper(like you did withMyLineMapper) instead of extending the entire reader. - Use custom
ItemProcessororItemWriterimplementations for business logic, not the core reader/writer infrastructure.
3. Optimize for Large Files
For 2.5GB CSV files, these tweaks will help prevent future issues:
- Split large files: Break the CSV into smaller, manageable chunks (e.g., 1GB each) and use
MultiResourceItemReaderto process them in parallel. - Tune chunk size: Adjust your
chunkSizeto balance performance and transaction stability. A range of 1000-5000 is typical for most databases—too small leads to excessive commits, too large risks timeouts. - Check JVM memory: Ensure your JVM has enough heap space (
-Xmxflag) to handle in-memory chunk processing without OOM errors.
Final Takeaway
Your core issue was almost certainly the non-thread-safe reader being used in a multi-threaded step. Spring Batch's default components handle these low-level details out of the box, so avoiding unnecessary customization of core infrastructure will save you a lot of headaches.
内容的提问来源于stack exchange,提问作者agl95

