Spring Batch中BeanIO解析嵌套XML的类型转换错误解决
问题背景
基于Spring Boot的Spring Batch应用处理XML输入文件时,执行importUserJob的step1步骤抛出类型转换异常:com.test.springbatchapp.model.ContactList无法转换为com.test.springbatchapp.model.Contact。
XML输入结构包含外层query节点,内部results容器下有多个重复的contact节点,当前BeanIO配置将results映射为ContactList对象,但Step期望处理的是单个Contact对象,导致类型不匹配。
相关代码与配置
XML输入文件
<?xml version="1.0" encoding="UTF-8" ?> <query> <id>123</id> <tracking>555</tracking> <results> <contact> <full_name> <first_name>John</first_name> <last_name>Doe</last_name> </full_name> <street>123 Main St</street> <city>Chicago</city> <state>IL</state> <zip>60610</zip> </contact> <contact> <full_name> <first_name>Jane</first_name> <last_name>Smith</last_name> </full_name> <street>123 Main St</street> <city>Miami</city> <state>FL</state> </contact> </results> </query>
原BeanIO配置(beanio-configuration.xml)
<beanio xmlns="http://www.beanio.org/2012/03"> <stream name="query" format="xml" strict="true" ignoreUnidentifiedRecords="true"> <record name="id"></record> <record name="tracking"></record> <record name="results" class="com.test.springbatchapp.model.ContactList"> <segment name="contact" collection="list" class="com.test.springbatchapp.model.Contact" occurs="0+"> <segment name="full_name"> <field name="firstName" xmlName="first_name" maxLength="20" /> <field name="lastName" xmlName="last_name" maxLength="30" /> </segment> <field name="street" maxLength="30" /> <field name="city" maxLength="25" /> <field name="state" minLength="2" maxLength="2" /> <field name="zip" regex="\d{5}" minOccurs="0" default="" /> </segment> </record> </stream> </beanio>
实体类
Contact.java:
@Data @AllArgsConstructor @NoArgsConstructor public class Contact { private String firstName; private String lastName; private String street; private String city; private String state; private String zip; }
ContactList.java:
@Data @AllArgsConstructor @NoArgsConstructor public class ContactList { private List<Contact> contact; }
原Reader配置
@Bean public ItemReader<Contact> reader() { BeanIOFlatFileItemReader<Contact> reader = new BeanIOFlatFileItemReader<>(); try { reader.setResource(new FileSystemResource(fileInputContact)); reader.setStreamName(inputContactStreamName); reader.setStreamMapping(new ClassPathResource(beanIoConfigurationXmlPath)); reader.setStreamFactory(StreamFactory.newInstance()); reader.getLineNumber(); reader.afterPropertiesSet(); } catch (Exception e) { log.error("ERROR: An issue occurred in the BeanIO Item Reader:: {} {}", e.getMessage(), e.getStackTrace()); } return reader; }
解决方案
方案一:修改BeanIO配置,直接返回单个Contact对象(推荐)
调整BeanIO配置,忽略外层无关节点,直接将每个contact节点映射为Contact对象,让Reader返回Step期望的类型:
修改后的beanio-configuration.xml:
<beanio xmlns="http://www.beanio.org/2012/03"> <stream name="query" format="xml" strict="true" ignoreUnidentifiedRecords="true"> <!-- 忽略外层query、id、tracking节点,直接读取results下的contact --> <record name="query" ignore="true"> <field name="id" ignore="true"/> <field name="tracking" ignore="true"/> <segment name="results" ignore="true"> <record name="contact" class="com.test.springbatchapp.model.Contact" occurs="0+"> <segment name="full_name"> <field name="firstName" xmlName="first_name" maxLength="20" /> <field name="lastName" xmlName="last_name" maxLength="30" /> </segment> <field name="street" maxLength="30" /> <field name="city" maxLength="25" /> <field name="state" minLength="2" maxLength="2" /> <field name="zip" regex="\d{5}" minOccurs="0" default="" /> </record> </segment> </record> </stream> </beanio>
说明:此配置会跳过query、id、tracking、results这些外层节点,直接识别并映射每个contact节点为Contact对象,Reader返回的类型与Step期望的Contact完全匹配,无需修改其他配置。
方案二:保留ContactList结构,通过Processor拆分列表
如果需要保留ContactList的映射结构,可调整Reader泛型,并通过ItemProcessor将ContactList拆分为单个Contact对象:
- 修改Reader泛型为ContactList:
@Bean public ItemReader<ContactList> reader() { BeanIOFlatFileItemReader<ContactList> reader = new BeanIOFlatFileItemReader<>(); try { reader.setResource(new FileSystemResource(fileInputContact)); reader.setStreamName(inputContactStreamName); reader.setStreamMapping(new ClassPathResource(beanIoConfigurationXmlPath)); reader.setStreamFactory(StreamFactory.newInstance()); reader.getLineNumber(); reader.afterPropertiesSet(); } catch (Exception e) { log.error("ERROR: An issue occurred in the BeanIO Item Reader:: {} {}", e.getMessage(), e.getStackTrace()); } return reader; }
- 修改Step配置与添加拆分Processor:
@Bean public Step step1(FlatFileItemWriter<Contact> writer) { return stepBuilderFactory.get("step1") .<ContactList, Contact> chunk(10) .reader(reader()) .processor(contactListSplitter()) .writer(writer) .build(); } @Bean public ItemProcessor<ContactList, List<Contact>> contactListSplitter() { // 返回ContactList中的contact列表,Spring Batch会自动展开列表元素到Chunk中 return contactList -> contactList.getContact(); }
说明:此方案下,Reader返回ContactList对象,Processor将其拆分为Contact列表,Spring Batch会自动将列表中的每个Contact作为独立元素传入Writer处理。
内容的提问来源于stack exchange,提问作者Somebody

