You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Spring Batch解析无分隔符定长文件及配置读取器

Spring Batch处理无分隔符定长文件的解决方案

嘿,这问题我熟!Spring Batch处理这种定长无分隔符的文件其实很直接,核心就是用它自带的FixedLengthTokenizer来按位置拆分字段,我一步步给你讲清楚。

1. 如何解析无分隔符的定长格式文件

定长文件的特点是每个字段的长度固定,靠起始/结束位置来区分,比如你给的示例行:120180208FAILED220180208SUCCES120170208SUCCES,其实是由3个相同结构的记录拼接而成,每个记录的结构是:

  • 编码:1位(比如第一个记录的1)
  • 日期:8位(20180208)
  • 状态:6位(FAILED)

解析这类文件的关键就是告诉Spring Batch每个字段的位置范围,它会自动帮你把字符串拆分成对应的字段值,之后再映射到你的实体类里。

2. 配置Spring Batch读取器(按起始/结束位置确定元素)

下面我给你两种常用的配置方式,Java配置(现在更主流)和XML配置,你可以按需选择。

先准备实体类

首先定义一个对应记录的实体类,用来存储解析后的数据:

public class TransactionRecord {
    private String code;
    private LocalDate date;
    private String status;

    // 必须提供无参构造器,BeanWrapperFieldSetMapper需要
    public TransactionRecord() {}

    // Getters and Setters
    public String getCode() { return code; }
    public void setCode(String code) { this.code = code; }
    public LocalDate getDate() { return date; }
    public void setDate(LocalDate date) { this.date = date; }
    public String getStatus() { return status; }
    public void setStatus(String status) { this.status = status; }
}

Java配置方式

这是现在Spring Boot项目里最常用的配置方式,用@Configuration类来定义读取器:

import org.springframework.batch.item.file.FlatFileItemReader;
import org.springframework.batch.item.file.builder.FlatFileItemReaderBuilder;
import org.springframework.batch.item.file.mapping.BeanWrapperFieldSetMapper;
import org.springframework.batch.item.file.mapping.FieldSetMapper;
import org.springframework.batch.item.file.transform.FixedLengthTokenizer;
import org.springframework.batch.item.file.transform.Range;
import org.springframework.context.annotation.Bean;
import org.springframework.context.annotation.Configuration;
import org.springframework.core.io.ClassPathResource;
import org.springframework.beans.propertyeditors.CustomDateEditor;

import java.time.LocalDate;
import java.time.format.DateTimeFormatter;
import java.util.Map;

@Configuration
public class BatchConfig {

    // 核心:定长文件读取器
    @Bean
    public FlatFileItemReader<TransactionRecord> fixedLengthTransactionReader() {
        return new FlatFileItemReaderBuilder<TransactionRecord>()
                .name("fixedLengthTransactionReader")
                .resource(new ClassPathResource("transactions.txt")) // 替换成你的文件路径
                .lineTokenizer(fixedLengthTokenizer())
                .fieldSetMapper(fieldSetMapper())
                .build();
    }

    // 配置定长字段拆分器:指定每个字段的位置范围
    private FixedLengthTokenizer fixedLengthTokenizer() {
        FixedLengthTokenizer tokenizer = new FixedLengthTokenizer();
        // 定义字段名称,要和实体类的属性名对应
        tokenizer.setNames("code", "date", "status");
        // 重点:Spring Batch的Range是从1开始计数的!
        tokenizer.setColumns(
                new Range(1, 1),    // 编码:第1位(长度1)
                new Range(2, 9),    // 日期:第2-9位(共8位)
                new Range(10, 15)   // 状态:第10-15位(共6位)
        );
        return tokenizer;
    }

    // 配置字段映射器:把拆分后的字段映射到实体类
    private FieldSetMapper<TransactionRecord> fieldSetMapper() {
        BeanWrapperFieldSetMapper<TransactionRecord> mapper = new BeanWrapperFieldSetMapper<>();
        mapper.setTargetType(TransactionRecord.class);
        // 自定义日期格式转换:把yyyyMMdd格式的字符串转成LocalDate
        mapper.setCustomEditors(Map.of(
                LocalDate.class, new CustomDateEditor(DateTimeFormatter.ofPattern("yyyyMMdd"), true)
        ));
        return mapper;
    }
}

处理一行多记录的情况

你的示例行里一行包含3个记录,上面的配置默认把整行当成一个记录解析,这显然不对。我们可以自定义LineMapper来拆分每行的多个记录:

// 在BatchConfig里添加这个Bean,替换之前的fixedLengthTransactionReader
@Bean
public FlatFileItemReader<TransactionRecord> multiRecordFixedLengthReader() {
    return new FlatFileItemReaderBuilder<TransactionRecord>()
            .name("multiRecordFixedLengthReader")
            .resource(new ClassPathResource("transactions.txt"))
            .lineMapper(multiRecordLineMapper())
            .build();
}

// 自定义LineMapper,拆分一行成多个记录
private LineMapper<TransactionRecord> multiRecordLineMapper() {
    return (line, lineNumber) -> {
        FixedLengthTokenizer tokenizer = fixedLengthTokenizer();
        FieldSetMapper<TransactionRecord> mapper = fieldSetMapper();
        int singleRecordLength = 15; // 单个记录的总长度:1+8+6=15
        List<TransactionRecord> records = new ArrayList<>();

        // 循环拆分每行的多个记录
        for (int i = 0; i < line.length(); i += singleRecordLength) {
            if (i + singleRecordLength > line.length()) {
                // 跳过不完整的记录,避免报错
                continue;
            }
            String singleRecordStr = line.substring(i, i + singleRecordLength);
            records.add(mapper.mapFieldSet(tokenizer.tokenize(singleRecordStr)));
        }

        // 如果需要返回多个记录,建议用CompositeItemReader或者自定义ItemReader
        // 这里为了演示,返回第一个记录,实际使用时可以调整逻辑返回所有
        return records.isEmpty() ? null : records.get(0);
    };
}

关键注意点

  • Range计数规则:Spring Batch的Range是从1开始的,不是0,这个很容易踩坑,一定要记住!
  • 日期格式化:如果字段是日期类型,必须配置自定义编辑器来转换格式,否则会抛出类型转换异常
  • 异常处理:可以给FlatFileItemReader配置skipPolicy来跳过格式错误的行,避免整个Job失败
  • 资源路径:如果文件在classpath下用ClassPathResource,如果是绝对路径用FileSystemResource

内容的提问来源于stack exchange,提问作者B.Nbl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:13:01