Java实现XML转集合遍历、按性别聚合学生数据及去重方案求解
问题根因
你现在的代码跑不出正确结果,是几个问题叠加导致的:
- XML文件本身格式不合法:示例XML没有包裹所有学生节点的外层根标签,第一个
<gender>M<gender>没有正确闭合,第一个学生节点错误使用自闭合标签<student/>,导致解析器无法识别全部3条学生数据。 - JAXB映射配置错误:代码里JAXBContext初始化传入的是
Company.class,和实际要解析的学生结构完全不匹配;同时Student实体类定义的firstName、lastName、contactNo字段在XML里根本不存在,根本读不到正确的属性值。 - 去重逻辑失效:
Stream.distinct()和HashSet去重依赖类的equals()和hashCode()方法,你的Student类没有重写这两个方法,默认继承Object类的实现是比较对象内存地址,不同new出来的对象地址永远不同,根本无法按字段值判断重复。 - 代码逻辑混乱:循环里往名为
student的集合塞拼接后的对象,最后去重操作却作用在另一个list集合上,还凭空出现Employee类型的列表定义,类型完全不匹配。
实现步骤
1. 修正XML格式
先把XML改成合法的、带根节点的结构,所有标签正确闭合:
<?xml version="1.0"?> <students> <student> <name>aaa</name> <gender>M</gender> <Address> <City>Auckland</City> <Zipcode>2310</Zipcode> </Address> </student> <student> <name>bbb</name> <gender>f</gender> <Address> <City>Wellington</City> <Zipcode>2310</Zipcode> </Address> </student> <student> <name>ccc</name> <gender>f</gender> <Address> <City>NorthIsland</City> <Zipcode>5671</Zipcode> </Address> </student> </students>
2. 编写正确的JAXB映射实体类
注意必须重写Student类的equals()和hashCode()方法,保证去重逻辑正常生效。如果不用Lombok,自己手写getter、setter、无参构造即可。
地址实体类:
import jakarta.xml.bind.annotation.XmlAccessType; import jakarta.xml.bind.annotation.XmlAccessorType; import jakarta.xml.bind.annotation.XmlRootElement; @XmlRootElement(name = "Address") @XmlAccessorType(XmlAccessType.FIELD) public class Address { private String City; private String Zipcode; // 自己补全getter、setter public String getCity() { return City; } public void setCity(String city) { City = city; } public String getZipcode() { return Zipcode; } public void setZipcode(String zipcode) { Zipcode = zipcode; } }
学生实体类:
import jakarta.xml.bind.annotation.XmlAccessType; import jakarta.xml.bind.annotation.XmlAccessorType; import jakarta.xml.bind.annotation.XmlRootElement; import java.util.Objects; @XmlRootElement(name = "student") @XmlAccessorType(XmlAccessType.FIELD) public class Student { private String name; private String gender; private Address Address; // 自己补全getter、setter public String getName() { return name; } public void setName(String name) { this.name = name; } public String getGender() { return gender; } public void setGender(String gender) { this.gender = gender; } public Address getAddress() { return Address; } public void setAddress(Address address) { Address = address; } // 必须重写equals和hashCode,按业务规则判断重复:这里按姓名+性别+地址判断为同一条记录 @Override public boolean equals(Object o) { if (this == o) return true; if (o == null || getClass() != o.getClass()) return false; Student student = (Student) o; return Objects.equals(name, student.name) && Objects.equals(gender, student.gender) && Objects.equals(Address, student.Address); } @Override public int hashCode() { return Objects.hash(name, gender, Address); } }
外层根节点实体类:
import jakarta.xml.bind.annotation.XmlAccessType; import jakarta.xml.bind.annotation.XmlAccessorType; import jakarta.xml.bind.annotation.XmlElement; import jakarta.xml.bind.annotation.XmlRootElement; import java.util.List; @XmlRootElement(name = "students") @XmlAccessorType(XmlAccessType.FIELD) public class Students { @XmlElement(name = "student") private List<Student> studentList; public List<Student> getStudentList() { return studentList; } public void setStudentList(List<Student> studentList) { this.studentList = studentList; } }
3. 实现解析、去重、遍历、聚合写文件逻辑
List本身是有序集合,直接通过get(index)方法就可以支持按索引遍历,不需要额外转换:
import jakarta.xml.bind.JAXBContext; import jakarta.xml.bind.Unmarshaller; import java.io.File; import java.io.FileWriter; import java.io.IOException; import java.util.List; import java.util.Map; import java.util.stream.Collectors; public class StudentXmlProcess { public static void main(String[] args) { try { // 解析XML为集合 File xmlFile = new File("/Downloads/student.xml"); JAXBContext jaxbContext = JAXBContext.newInstance(Students.class); Unmarshaller unmarshaller = jaxbContext.createUnmarshaller(); Students root = (Students) unmarshaller.unmarshal(xmlFile); List<Student> originList = root.getStudentList(); // 去重 List<Student> distinctList = originList.stream().distinct().collect(Collectors.toList()); // 按索引遍历 System.out.println("=== 按索引遍历去重后的学生数据 ==="); for (int i = 0; i < distinctList.size(); i++) { Student stu = distinctList.get(i); System.out.printf("索引:%d,姓名:%s,性别:%s,城市:%s%n", i, stu.getName(), stu.getGender(), stu.getAddress().getCity()); } // 按性别聚合,统一性别大小写避免M/m、F/f被分到不同组 Map<String, List<Student>> genderGroupMap = distinctList.stream() .collect(Collectors.groupingBy(stu -> stu.getGender().toUpperCase())); // 不同性别写入不同文件 for (Map.Entry<String, List<Student>> entry : genderGroupMap.entrySet()) { String gender = entry.getKey(); List<Student> groupStudents = entry.getValue(); String outputPath = "/Downloads/student_gender_" + gender + ".txt"; try (FileWriter writer = new FileWriter(outputPath)) { for (Student stu : groupStudents) { writer.write(String.format("姓名:%s,城市:%s,邮编:%s%n", stu.getName(), stu.getAddress().getCity(), stu.getAddress().getZipcode())); } } catch (IOException e) { e.printStackTrace(); } System.out.printf("性别%s共%d条数据,已写入文件:%s%n", gender, groupStudents.size(), outputPath); } } catch (Exception e) { e.printStackTrace(); } } }
注意事项
- JDK8自带JAXB包,JDK9及以上版本需要单独引入JAXB依赖,否则会报类找不到错误。
- 如果去重规则不需要全字段匹配,比如只要姓名和性别相同就算重复,直接修改
equals()和hashCode()方法里的判断字段即可。 - 聚合前统一性别字段的大小写,避免因为大小写格式不统一导致分组错误。
内容的提问来源于stack exchange,提问作者newtoJava
相关产品推荐
相关产品推荐

