Spring Batch解析定长Flat XML文件的优化方案咨询
我有如下格式的定长Flat XML文件:
<?xml version="1.0" encoding="UTF-8"?> <File fileId="123" xmlns="abc:XYZ" > ABC123411/10/20 XBC128911/10/20 BCD456711/23/22 </File>
需要将其中每行定长内容解析为Content对象,示例映射关系为:
ABC123411/10/20 → name: ABC, id: 1234, Date: 11/10/20
Content类定义如下:
public class Content { private id; private name; private date; // getters }
目前我使用StaxEventItemReader结合Jaxb2Marshaller的配置,只能将<File>节点内的文本内容整体读取为TestRecord的data属性,后续需额外编写映射器拆分字符串并转换为Content对象。想咨询是否可通过Spring Batch的XML配置直接实现解析,无需额外映射逻辑?
可以通过Spring Batch现有组件的组合配置实现,无需额外编写独立映射器,核心是把XML节点内的文本提取后交给定长文件阅读器处理:
第一步:用StaxEventItemReader提取File节点内的文本
先定义一个承载File节点文本的类,通过JAXB注解绑定XML结构:@XmlRootElement(name = "File", namespace = "abc:XYZ") public class FileContentRecord { @XmlValue private String content; // getter、setter方法 }配置StaxEventItemReader读取该节点,将内部所有文本作为一个字符串返回:
<bean id="staxReader" class="org.springframework.batch.item.xml.StaxEventItemReader"> <property name="resource" value="classpath:your-input.xml"/> <property name="unmarshaller" ref="jaxbMarshaller"/> <property name="fragmentRootElementName" value="File"/> </bean> <bean id="jaxbMarshaller" class="org.springframework.oxm.jaxb.Jaxb2Marshaller"> <property name="classesToBeBound" value="com.example.FileContentRecord"/> <property name="contextPaths" value="abc:XYZ"/> </bean>第二步:拆分文本为单行列表
编写简单处理器,把整段文本按换行符拆分:public class ContentSplitterProcessor implements ItemProcessor<FileContentRecord, List<String>> { @Override public List<String> process(FileContentRecord item) throws Exception { String rawContent = item.getContent().trim(); return Arrays.asList(rawContent.split("\\r?\\n")); } }第三步:用FlatFileItemReader处理定长行
配置定长阅读器,直接将每行映射为Content对象:<bean id="fixedLengthReader" class="org.springframework.batch.item.file.FlatFileItemReader"> <!-- 用虚拟资源占位,实际内容来自上游拆分结果 --> <property name="resource" value="classpath:dummy.txt"/> <property name="lineMapper"> <bean class="org.springframework.batch.item.file.mapping.DefaultLineMapper"> <property name="fieldSetMapper"> <bean class="org.springframework.batch.item.file.mapping.BeanWrapperFieldSetMapper"> <property name="targetType" value="com.example.Content"/> <property name="customEditors"> <map> <entry key="java.util.Date"> <bean class="org.springframework.batch.item.file.transform.DateFieldEditor"> <property name="pattern" value="MM/dd/yy"/> </bean> </entry> </map> </property> </bean> </property> <property name="lineTokenizer"> <bean class="org.springframework.batch.item.file.transform.FixedLengthTokenizer"> <!-- 对应name(3位)、id(4位)、date(8位)的定长位置 --> <property name="columns" value="1-3, 4-7, 8-15"/> <property name="names" value="name, id, date"/> </bean> </property> </bean> </property> </bean>第四步:串联组件到作业流
在作业的step中,先通过staxReader读取File节点文本,经处理器拆分为行列表,再遍历列表调用fixedLengthReader处理每行,最终得到Content对象。可以通过自定义ChunkProcessor或者使用CompositeItemWriter来实现行的遍历转换。
这种方式完全基于Spring Batch的标准组件配置,不需要额外的独立映射逻辑,直接完成从XML内定长文本到Content对象的解析。
内容的提问来源于stack exchange,提问作者CuriousToLearn

