You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spring Batch定长FlatFileItemReader校验尾行存在否则任务失败实现方案

Spring Batch 大文件分块处理的尾行校验方案

针对定长记录大文件带固定尾行的校验需求,不要用全量加载文件的方式实现,会完全抵消分块处理的内存优势,以下是可直接落地的低内存占用实现:

核心逻辑

  • Step启动前用随机文件流定位到文件尾部,仅读取最后几十到上百字节内容校验尾行格式,不加载全量文件,内存开销和文件大小完全无关
  • 改造原有定长读取器,识别到尾行时直接终止读取,避免尾行被解析为业务记录
  • (可选)增加记录数双重校验:比对尾行标注的总记录数和实际读取的业务记录数,覆盖文件中途截断的极端场景

代码实现

1. 尾行校验监听器

这个监听器会在Step执行前自动完成尾行校验,校验失败直接抛出异常终止作业,校验通过后才会启动分块读流程:

@Component
@StepScope
public class FileTailCheckListener implements StepExecutionListener {
    // 替换为你实际的尾行固定前缀
    private static final String TAIL_PREFIX = "Total number of records: ";
    // 尾行最大预估长度,留足冗余即可,比如存10位数字加前缀总长度不超过50,设100足够
    private static final int TAIL_BUFFER_LENGTH = 100;

    @Value("#{jobParameters['inputFile']}")
    private Resource inputResource;

    @Override
    public void beforeStep(StepExecution stepExecution) {
        try (RandomAccessFile raf = new RandomAccessFile(inputResource.getFile(), "r")) {
            long fileLen = raf.length();
            // 定位到文件尾部往前偏移的位置,兼容文件总长度小于缓冲长度的情况
            long seekPos = Math.max(0, fileLen - TAIL_BUFFER_LENGTH);
            raf.seek(seekPos);

            String lastLine = null;
            String currentLine;
            // 如果不是从文件头开始读,第一行大概率是被截断的残行,直接丢弃
            if (seekPos > 0) {
                raf.readLine();
            }
            // 遍历读到文件末尾,取最后一行
            while ((currentLine = raf.readLine()) != null) {
                lastLine = currentLine;
            }

            // 尾行格式校验,不通过直接抛异常让作业失败
            if (lastLine == null || !lastLine.startsWith(TAIL_PREFIX)) {
                throw new IllegalStateException("文件校验失败:未找到符合格式的尾行,文件不完整");
            }

            // 可选:解析尾行中的总记录数,存入上下文供后续比对
            int expectedCount = Integer.parseInt(lastLine.substring(TAIL_PREFIX.length()).trim());
            stepExecution.getExecutionContext().putInt("expectedTotal", expectedCount);
        } catch (IOException e) {
            throw new RuntimeException("文件读取失败,无法完成尾行校验", e);
        }
    }

    @Override
    public ExitStatus afterStep(StepExecution stepExecution) {
        // 可选:比对实际读取记录数和尾行标注的记录数
        int expectedTotal = stepExecution.getExecutionContext().getInt("expectedTotal", -1);
        int readCount = stepExecution.getReadCount();
        if (expectedTotal != -1 && expectedTotal != readCount) {
            throw new IllegalStateException(String.format("文件记录数不匹配:尾行标注共%d条,实际读取%d条,文件可能截断", expectedTotal, readCount));
        }
        return stepExecution.getExitStatus();
    }
}

2. 支持尾行识别的定长读取器

继承原生FlatFileItemReader,重写行读取逻辑,读到尾行时直接返回终止标识,避免尾行进入定长解析流程报错,也不会被当成业务数据处理:

public class TailSkipFlatFileReader<T> extends FlatFileItemReader<T> {
    private static final String TAIL_PREFIX = "Total number of records: ";

    @Override
    public String readLine() throws Exception {
        String line = super.readLine();
        // 读到尾行直接返回null,告知框架业务数据已全部读取完成
        if (line != null && line.startsWith(TAIL_PREFIX)) {
            return null;
        }
        return line;
    }
}

3. 修改原有作业配置

把自定义读取器和校验监听器注入到Step配置中即可:

@Configuration
public class JobConfig {

    @Bean
    public Step exampleLoad(
            StepBuilderFactory stepBuilderFactory,
            FileTailCheckListener tailCheckListener) {
        return stepBuilderFactory.get("exampleLoad")
                .<ExampleRecord, ExampleEntity>chunk(5000)
                .reader(exampleReader())
                .processor(exampleProcessor())
                .writer(exampleWriter())
                // 注册尾行校验监听器
                .listener(tailCheckListener)
                .build();
    }

    @Bean
    @StepScope
    public FlatFileItemReader<ExampleRecord> exampleReader() {
        // 替换为自定义的尾行跳过读取器
        return new TailSkipFlatFileReader<ExampleRecord>()
                .name("exampleReader")
                .resource(...)
                .fixedLength()
                .strict(false)
                .columns(new Range(1, 2), new Range(3, 4), new Range(5, 6))
                .names("a", "b", "c")
                .targetType(ExampleRecord.class)
                .build();
    }

    // processor/writer 保持原有逻辑不变
}

注意事项

  • 尾行校验环节的内存占用固定在百字节级别,哪怕是几十GB的大文件也不会出现OOM,完全适配分块处理场景
  • 若文件使用UTF-8、GBK等多字节编码,随机流定位时可能截断半个字符,因为我们会丢弃定位后的第一行残行,完全不会影响最后一行的读取正确性
  • 建议保留记录数比对逻辑,可以覆盖“尾行存在但文件传输中途丢失内容”的极端异常场景
  • 若尾行格式有变化,只需要修改TAIL_PREFIX常量和对应的解析逻辑即可,不需要调整核心流程

内容的提问来源于stack exchange,提问作者Tom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 10:51:22