Spring Batch定长FlatFileItemReader校验尾行存在否则任务失败实现方案
Spring Batch 大文件分块处理的尾行校验方案
针对定长记录大文件带固定尾行的校验需求,不要用全量加载文件的方式实现,会完全抵消分块处理的内存优势,以下是可直接落地的低内存占用实现:
核心逻辑
- Step启动前用随机文件流定位到文件尾部,仅读取最后几十到上百字节内容校验尾行格式,不加载全量文件,内存开销和文件大小完全无关
- 改造原有定长读取器,识别到尾行时直接终止读取,避免尾行被解析为业务记录
- (可选)增加记录数双重校验:比对尾行标注的总记录数和实际读取的业务记录数,覆盖文件中途截断的极端场景
代码实现
1. 尾行校验监听器
这个监听器会在Step执行前自动完成尾行校验,校验失败直接抛出异常终止作业,校验通过后才会启动分块读流程:
@Component @StepScope public class FileTailCheckListener implements StepExecutionListener { // 替换为你实际的尾行固定前缀 private static final String TAIL_PREFIX = "Total number of records: "; // 尾行最大预估长度,留足冗余即可,比如存10位数字加前缀总长度不超过50,设100足够 private static final int TAIL_BUFFER_LENGTH = 100; @Value("#{jobParameters['inputFile']}") private Resource inputResource; @Override public void beforeStep(StepExecution stepExecution) { try (RandomAccessFile raf = new RandomAccessFile(inputResource.getFile(), "r")) { long fileLen = raf.length(); // 定位到文件尾部往前偏移的位置,兼容文件总长度小于缓冲长度的情况 long seekPos = Math.max(0, fileLen - TAIL_BUFFER_LENGTH); raf.seek(seekPos); String lastLine = null; String currentLine; // 如果不是从文件头开始读,第一行大概率是被截断的残行,直接丢弃 if (seekPos > 0) { raf.readLine(); } // 遍历读到文件末尾,取最后一行 while ((currentLine = raf.readLine()) != null) { lastLine = currentLine; } // 尾行格式校验,不通过直接抛异常让作业失败 if (lastLine == null || !lastLine.startsWith(TAIL_PREFIX)) { throw new IllegalStateException("文件校验失败:未找到符合格式的尾行,文件不完整"); } // 可选:解析尾行中的总记录数,存入上下文供后续比对 int expectedCount = Integer.parseInt(lastLine.substring(TAIL_PREFIX.length()).trim()); stepExecution.getExecutionContext().putInt("expectedTotal", expectedCount); } catch (IOException e) { throw new RuntimeException("文件读取失败,无法完成尾行校验", e); } } @Override public ExitStatus afterStep(StepExecution stepExecution) { // 可选:比对实际读取记录数和尾行标注的记录数 int expectedTotal = stepExecution.getExecutionContext().getInt("expectedTotal", -1); int readCount = stepExecution.getReadCount(); if (expectedTotal != -1 && expectedTotal != readCount) { throw new IllegalStateException(String.format("文件记录数不匹配:尾行标注共%d条,实际读取%d条,文件可能截断", expectedTotal, readCount)); } return stepExecution.getExitStatus(); } }
2. 支持尾行识别的定长读取器
继承原生FlatFileItemReader,重写行读取逻辑,读到尾行时直接返回终止标识,避免尾行进入定长解析流程报错,也不会被当成业务数据处理:
public class TailSkipFlatFileReader<T> extends FlatFileItemReader<T> { private static final String TAIL_PREFIX = "Total number of records: "; @Override public String readLine() throws Exception { String line = super.readLine(); // 读到尾行直接返回null,告知框架业务数据已全部读取完成 if (line != null && line.startsWith(TAIL_PREFIX)) { return null; } return line; } }
3. 修改原有作业配置
把自定义读取器和校验监听器注入到Step配置中即可:
@Configuration public class JobConfig { @Bean public Step exampleLoad( StepBuilderFactory stepBuilderFactory, FileTailCheckListener tailCheckListener) { return stepBuilderFactory.get("exampleLoad") .<ExampleRecord, ExampleEntity>chunk(5000) .reader(exampleReader()) .processor(exampleProcessor()) .writer(exampleWriter()) // 注册尾行校验监听器 .listener(tailCheckListener) .build(); } @Bean @StepScope public FlatFileItemReader<ExampleRecord> exampleReader() { // 替换为自定义的尾行跳过读取器 return new TailSkipFlatFileReader<ExampleRecord>() .name("exampleReader") .resource(...) .fixedLength() .strict(false) .columns(new Range(1, 2), new Range(3, 4), new Range(5, 6)) .names("a", "b", "c") .targetType(ExampleRecord.class) .build(); } // processor/writer 保持原有逻辑不变 }
注意事项
- 尾行校验环节的内存占用固定在百字节级别,哪怕是几十GB的大文件也不会出现OOM,完全适配分块处理场景
- 若文件使用UTF-8、GBK等多字节编码,随机流定位时可能截断半个字符,因为我们会丢弃定位后的第一行残行,完全不会影响最后一行的读取正确性
- 建议保留记录数比对逻辑,可以覆盖“尾行存在但文件传输中途丢失内容”的极端异常场景
- 若尾行格式有变化,只需要修改
TAIL_PREFIX常量和对应的解析逻辑即可,不需要调整核心流程
内容的提问来源于stack exchange,提问作者Tom
相关产品推荐
相关产品推荐

