如何使用Spring Batch解析无分隔符定长文件及配置读取器
Spring Batch处理无分隔符定长文件的解决方案
嘿,这问题我熟!Spring Batch处理这种定长无分隔符的文件其实很直接,核心就是用它自带的FixedLengthTokenizer来按位置拆分字段,我一步步给你讲清楚。
1. 如何解析无分隔符的定长格式文件
定长文件的特点是每个字段的长度固定,靠起始/结束位置来区分,比如你给的示例行:120180208FAILED220180208SUCCES120170208SUCCES,其实是由3个相同结构的记录拼接而成,每个记录的结构是:
- 编码:1位(比如第一个记录的
1) - 日期:8位(
20180208) - 状态:6位(
FAILED)
解析这类文件的关键就是告诉Spring Batch每个字段的位置范围,它会自动帮你把字符串拆分成对应的字段值,之后再映射到你的实体类里。
2. 配置Spring Batch读取器(按起始/结束位置确定元素)
下面我给你两种常用的配置方式,Java配置(现在更主流)和XML配置,你可以按需选择。
先准备实体类
首先定义一个对应记录的实体类,用来存储解析后的数据:
public class TransactionRecord { private String code; private LocalDate date; private String status; // 必须提供无参构造器,BeanWrapperFieldSetMapper需要 public TransactionRecord() {} // Getters and Setters public String getCode() { return code; } public void setCode(String code) { this.code = code; } public LocalDate getDate() { return date; } public void setDate(LocalDate date) { this.date = date; } public String getStatus() { return status; } public void setStatus(String status) { this.status = status; } }
Java配置方式
这是现在Spring Boot项目里最常用的配置方式,用@Configuration类来定义读取器:
import org.springframework.batch.item.file.FlatFileItemReader; import org.springframework.batch.item.file.builder.FlatFileItemReaderBuilder; import org.springframework.batch.item.file.mapping.BeanWrapperFieldSetMapper; import org.springframework.batch.item.file.mapping.FieldSetMapper; import org.springframework.batch.item.file.transform.FixedLengthTokenizer; import org.springframework.batch.item.file.transform.Range; import org.springframework.context.annotation.Bean; import org.springframework.context.annotation.Configuration; import org.springframework.core.io.ClassPathResource; import org.springframework.beans.propertyeditors.CustomDateEditor; import java.time.LocalDate; import java.time.format.DateTimeFormatter; import java.util.Map; @Configuration public class BatchConfig { // 核心:定长文件读取器 @Bean public FlatFileItemReader<TransactionRecord> fixedLengthTransactionReader() { return new FlatFileItemReaderBuilder<TransactionRecord>() .name("fixedLengthTransactionReader") .resource(new ClassPathResource("transactions.txt")) // 替换成你的文件路径 .lineTokenizer(fixedLengthTokenizer()) .fieldSetMapper(fieldSetMapper()) .build(); } // 配置定长字段拆分器:指定每个字段的位置范围 private FixedLengthTokenizer fixedLengthTokenizer() { FixedLengthTokenizer tokenizer = new FixedLengthTokenizer(); // 定义字段名称,要和实体类的属性名对应 tokenizer.setNames("code", "date", "status"); // 重点:Spring Batch的Range是从1开始计数的! tokenizer.setColumns( new Range(1, 1), // 编码:第1位(长度1) new Range(2, 9), // 日期:第2-9位(共8位) new Range(10, 15) // 状态:第10-15位(共6位) ); return tokenizer; } // 配置字段映射器:把拆分后的字段映射到实体类 private FieldSetMapper<TransactionRecord> fieldSetMapper() { BeanWrapperFieldSetMapper<TransactionRecord> mapper = new BeanWrapperFieldSetMapper<>(); mapper.setTargetType(TransactionRecord.class); // 自定义日期格式转换:把yyyyMMdd格式的字符串转成LocalDate mapper.setCustomEditors(Map.of( LocalDate.class, new CustomDateEditor(DateTimeFormatter.ofPattern("yyyyMMdd"), true) )); return mapper; } }
处理一行多记录的情况
你的示例行里一行包含3个记录,上面的配置默认把整行当成一个记录解析,这显然不对。我们可以自定义LineMapper来拆分每行的多个记录:
// 在BatchConfig里添加这个Bean,替换之前的fixedLengthTransactionReader @Bean public FlatFileItemReader<TransactionRecord> multiRecordFixedLengthReader() { return new FlatFileItemReaderBuilder<TransactionRecord>() .name("multiRecordFixedLengthReader") .resource(new ClassPathResource("transactions.txt")) .lineMapper(multiRecordLineMapper()) .build(); } // 自定义LineMapper,拆分一行成多个记录 private LineMapper<TransactionRecord> multiRecordLineMapper() { return (line, lineNumber) -> { FixedLengthTokenizer tokenizer = fixedLengthTokenizer(); FieldSetMapper<TransactionRecord> mapper = fieldSetMapper(); int singleRecordLength = 15; // 单个记录的总长度:1+8+6=15 List<TransactionRecord> records = new ArrayList<>(); // 循环拆分每行的多个记录 for (int i = 0; i < line.length(); i += singleRecordLength) { if (i + singleRecordLength > line.length()) { // 跳过不完整的记录,避免报错 continue; } String singleRecordStr = line.substring(i, i + singleRecordLength); records.add(mapper.mapFieldSet(tokenizer.tokenize(singleRecordStr))); } // 如果需要返回多个记录,建议用CompositeItemReader或者自定义ItemReader // 这里为了演示,返回第一个记录,实际使用时可以调整逻辑返回所有 return records.isEmpty() ? null : records.get(0); }; }
关键注意点
- Range计数规则:Spring Batch的
Range是从1开始的,不是0,这个很容易踩坑,一定要记住! - 日期格式化:如果字段是日期类型,必须配置自定义编辑器来转换格式,否则会抛出类型转换异常
- 异常处理:可以给
FlatFileItemReader配置skipPolicy来跳过格式错误的行,避免整个Job失败 - 资源路径:如果文件在classpath下用
ClassPathResource,如果是绝对路径用FileSystemResource
内容的提问来源于stack exchange,提问作者B.Nbl
相关产品推荐
相关产品推荐

