ElementTree解析XML:按规则拼接property文本并去重重复父节点
解决方案
嘿,你的现有思路得调整一下哦——当前代码把所有property的文本混进一个大列表后去重,完全破坏了每个parent的分组结构,既没法区分不同父节点的文本集合,也实现不了你要的“每个properties层级最后一个文本用&替代!”的需求。咱们得换个思路,先按父节点分组处理,再完成去重和拼接:
第一步:按父节点提取文本列表
首先得遍历每个parent节点,把每个父节点下的property文本单独存成一个列表,这样才能保证每个父节点的独立结构:
import xml.etree.ElementTree as ET # 假设你已经解析好XML拿到了root节点 # tree = ET.parse('your_file.xml') # root = tree.getroot() parent_text_groups = [] # 匹配所有以parent-开头的节点 for parent in root.findall('.//*[starts-with(local-name(), "parent-")]'): # 提取当前父节点下所有property的文本 prop_texts = [prop.text for prop in parent.find('./properties').findall('property')] parent_text_groups.append(prop_texts)
第二步:剔除重复的父节点文本列表
要保留第一次出现的列表,后续完全重复的直接丢弃。这里用集合来记录已经见过的列表(注意列表不能直接存集合,得转成不可变的元组):
unique_groups = [] seen_groups = set() for group in parent_text_groups: group_tuple = tuple(group) if group_tuple not in seen_groups: seen_groups.add(group_tuple) unique_groups.append(group)
第三步:完成拼接逻辑
现在对每个唯一的文本列表,先内部用!连接,然后把所有这些连接后的结果用&串起来——这正好符合你要的“每个properties层级最后一个文本用&替代!”的效果(每个父节点内部用!连,不同父节点之间用&分隔,等价于每个父节点的最后一个元素后用&衔接下一个父节点的内容):
final_result = '&'.join(['!'.join(group) for group in unique_groups])
完整可运行示例
把上面的步骤整合起来,用你的示例XML测试:
import xml.etree.ElementTree as ET # 你的示例XML内容 sample_xml = '''<root> <parent-1> <text>blah-1</text> <properties> <property type="R" id="0005">text-value-A</property> <property type="W" id="0003">text-value-B</property> <property type="H" id="0002">text-value-C</property> <property type="W" id="0008">text-value-D</property> </properties> </parent-1> <parent-2> <text>blah-2</text> <properties> <property type="W" id="0004">text-value-A</property> <property type="H" id="0087">text-value-B</property> </properties> </parent-2> <parent-3> <text>blah-3</text> <properties> <property type="H" id="0087">text-value-C</property> <property type="R" id="0008">text-value-A</property> </properties> </parent-3> <parent-4> <text>blah-4</text> <properties> <property type="H" id="0019">text-value-C</property> <property type="R" id="0060">text-value-A</property> </properties> </parent-4> </root>''' # 解析XML root = ET.fromstring(sample_xml) # 提取父节点文本分组 parent_text_groups = [] for parent in root.findall('.//*[starts-with(local-name(), "parent-")]'): prop_texts = [prop.text for prop in parent.find('./properties').findall('property')] parent_text_groups.append(prop_texts) # 去重 unique_groups = [] seen_groups = set() for group in parent_text_groups: group_tuple = tuple(group) if group_tuple not in seen_groups: seen_groups.add(group_tuple) unique_groups.append(group) # 拼接结果 final_result = '&'.join(['!'.join(group) for group in unique_groups]) print(final_result) # 输出正好是你要的:text-value-A!text-value-B!text-value-C!text-value-D&text-value-A!text-value-B&text-value-C!text-value-A
为啥你的原思路不行?
你的代码把所有property的文本混进一个列表后去重,不仅丢失了每个父节点的分组信息,还会错误地去掉单个重复的文本(比如如果两个父节点都有text-value-A,原代码会只保留一个),这完全不符合你“剔除整个文本列表重复的父节点”的需求。必须先按父节点分组,再处理去重和拼接,才能达到预期效果。
内容的提问来源于stack exchange,提问作者bt123
相关产品推荐
相关产品推荐

