Spring Batch处理CSV文件遇IncorrectTokenCountException如何跳过错误行
Great question! This is a super common scenario when dealing with messy CSV files in Spring Batch. There are two straightforward approaches to handle this, depending on whether you want to catch the exception directly or pre-validate rows to avoid errors altogether.
Approach 1: Skip Rows via Exception Handling (Recommended)
The simplest way is to configure your batch step to catch the IncorrectTokenCountException thrown by the reader, then skip the problematic row and keep processing. Here's how to implement this with both Java and XML configuration:
Java Configuration
@Configuration @EnableBatchProcessing public class BatchConfig { private final JobBuilderFactory jobBuilderFactory; private final StepBuilderFactory stepBuilderFactory; // Constructor injection for factories public BatchConfig(JobBuilderFactory jobBuilderFactory, StepBuilderFactory stepBuilderFactory) { this.jobBuilderFactory = jobBuilderFactory; this.stepBuilderFactory = stepBuilderFactory; } @Bean public FlatFileItemReader<MyDto> csvReader() { return new FlatFileItemReaderBuilder<MyDto>() .name("csvReader") .resource(new ClassPathResource("data.csv")) .delimited() .names("field1", "field2", "field3", "field4", "field5") // Your 5 target fields .fieldSetMapper(new BeanWrapperFieldSetMapper<>() {{ setTargetType(MyDto.class); }}) .build(); } @Bean public Step processingStep(ItemReader<MyDto> reader, ItemWriter<MyDto> writer) { return stepBuilderFactory.get("processingStep") .<MyDto, MyDto>chunk(10) .reader(reader) .writer(writer) // Skip rows that trigger the token count exception .skip(IncorrectTokenCountException.class) // Set a limit to avoid infinite loops if all rows are invalid .skipLimit(100) // Optional: Add a listener to log skipped rows for debugging .listener(new SkipListener<MyDto, MyDto>() { private final Logger log = LoggerFactory.getLogger(getClass()); @Override public void onSkipInRead(Throwable t) { if (t instanceof IncorrectTokenCountException) { IncorrectTokenCountException ex = (IncorrectTokenCountException) t; log.warn("Skipping invalid row: {}", ex.getInput()); } } }) .build(); } // Define your job, writer, and other components here }
XML Configuration
If you're using XML-based setup:
<beans xmlns="http://www.springframework.org/schema/beans" xmlns:batch="http://www.springframework.org/schema/batch" xsi:schemaLocation="http://www.springframework.org/schema/beans http://www.springframework.org/schema/beans/spring-beans.xsd http://www.springframework.org/schema/batch http://www.springframework.org/schema/batch/spring-batch.xsd"> <batch:job id="dataProcessingJob"> <batch:step id="processingStep"> <batch:tasklet> <batch:chunk reader="csvReader" writer="csvWriter" commit-interval="10"> <batch:skip-policy> <bean class="org.springframework.batch.core.step.skip.SimpleSkipPolicy"> <property name="skipLimit" value="100"/> <property name="skippableExceptions"> <map> <entry key="org.springframework.batch.item.file.transform.IncorrectTokenCountException" value="true"/> </map> </property> </bean> </batch:skip-policy> </batch:chunk> <batch:listeners> <batch:listener ref="skipLoggerListener"/> </batch:listeners> </batch:tasklet> </batch:step> </batch:job> <bean id="csvReader" class="org.springframework.batch.item.file.FlatFileItemReader"> <property name="resource" value="classpath:data.csv"/> <property name="lineMapper"> <bean class="org.springframework.batch.item.file.mapping.DefaultLineMapper"> <property name="lineTokenizer"> <bean class="org.springframework.batch.item.file.transform.DelimitedLineTokenizer"> <property name="names" value="field1,field2,field3,field4,field5"/> </bean> </property> <property name="fieldSetMapper"> <bean class="org.springframework.batch.item.file.mapping.BeanWrapperFieldSetMapper"> <property name="targetType" value="com.example.MyDto"/> </bean> </property> </bean> </property> </bean> <bean id="skipLoggerListener" class="com.example.SkipLoggerListener"/> <!-- Define your writer and other beans here --> </beans>
Skip Listener Class
public class SkipLoggerListener implements SkipListener<MyDto, MyDto> { private static final Logger log = LoggerFactory.getLogger(SkipLoggerListener.class); @Override public void onSkipInRead(Throwable t) { if (t instanceof IncorrectTokenCountException) { IncorrectTokenCountException ex = (IncorrectTokenCountException) t; log.warn("Skipping row with incorrect token count: {}", ex.getInput()); } } // Empty implementations for unused methods @Override public void onSkipInWrite(MyDto item, Throwable t) {} @Override public void onSkipInProcess(MyDto item, Throwable t) {} }
Approach 2: Pre-Validate Rows (Avoid Exceptions)
If you prefer to avoid throwing exceptions entirely, you can configure the tokenizer to accept variable token counts, then validate rows in a processor and skip those with missing fields:
Step 1: Configure Tokenizer for Non-Strict Mode
@Bean public FlatFileItemReader<MyDto> csvReader() { DelimitedLineTokenizer tokenizer = new DelimitedLineTokenizer(); tokenizer.setNames("field1", "field2", "field3", "field4", "field5"); tokenizer.setStrict(false); // Disable strict token count checking DefaultLineMapper<MyDto> lineMapper = new DefaultLineMapper<>(); lineMapper.setLineTokenizer(tokenizer); lineMapper.setFieldSetMapper(new BeanWrapperFieldSetMapper<>() {{ setTargetType(MyDto.class); }}); return new FlatFileItemReaderBuilder<MyDto>() .name("csvReader") .resource(new ClassPathResource("data.csv")) .lineMapper(lineMapper) .build(); }
Step 2: Add Validation Processor
@Bean public ItemProcessor<MyDto, MyDto> validationProcessor() { return item -> { // Check if all required fields are present (adjust logic to match your DTO) if (item.getField1() == null || item.getField2() == null || item.getField3() == null || item.getField4() == null || item.getField5() == null) { log.warn("Skipping row with missing fields: {}", item); return null; // Return null to mark item for skipping } return item; }; }
Step 3: Configure Step to Skip Null Items
@Bean public Step processingStep(ItemReader<MyDto> reader, ItemProcessor<MyDto, MyDto> processor, ItemWriter<MyDto> writer) { return stepBuilderFactory.get("processingStep") .<MyDto, MyDto>chunk(10) .reader(reader) .processor(processor) .writer(writer) .skipNullItems(true) // Skip items returned as null from the processor .skipLimit(100) .build(); }
Key Tips
- Skip Limit: Always set a
skipLimitto prevent your job from running indefinitely if every row is invalid. - Logging: Use a
SkipListenerto track skipped rows – this is critical for debugging and auditing. - Choose Your Approach: Approach 1 is ideal for quick fixes when you just need to skip rows with wrong token counts. Approach 2 gives you more control over validation logic (e.g., skip only if specific fields are missing).
内容的提问来源于stack exchange,提问作者Jacel

