You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ElementTree解析XML:按规则拼接property文本并去重重复父节点

解决方案

嘿,你的现有思路得调整一下哦——当前代码把所有property的文本混进一个大列表后去重,完全破坏了每个parent的分组结构,既没法区分不同父节点的文本集合,也实现不了你要的“每个properties层级最后一个文本用&替代!”的需求。咱们得换个思路,先按父节点分组处理,再完成去重和拼接:

第一步:按父节点提取文本列表

首先得遍历每个parent节点,把每个父节点下的property文本单独存成一个列表,这样才能保证每个父节点的独立结构:

import xml.etree.ElementTree as ET

# 假设你已经解析好XML拿到了root节点
# tree = ET.parse('your_file.xml')
# root = tree.getroot()

parent_text_groups = []
# 匹配所有以parent-开头的节点
for parent in root.findall('.//*[starts-with(local-name(), "parent-")]'):
    # 提取当前父节点下所有property的文本
    prop_texts = [prop.text for prop in parent.find('./properties').findall('property')]
    parent_text_groups.append(prop_texts)

第二步:剔除重复的父节点文本列表

要保留第一次出现的列表,后续完全重复的直接丢弃。这里用集合来记录已经见过的列表(注意列表不能直接存集合,得转成不可变的元组):

unique_groups = []
seen_groups = set()
for group in parent_text_groups:
    group_tuple = tuple(group)
    if group_tuple not in seen_groups:
        seen_groups.add(group_tuple)
        unique_groups.append(group)

第三步:完成拼接逻辑

现在对每个唯一的文本列表,先内部用!连接,然后把所有这些连接后的结果用&串起来——这正好符合你要的“每个properties层级最后一个文本用&替代!”的效果(每个父节点内部用!连,不同父节点之间用&分隔,等价于每个父节点的最后一个元素后用&衔接下一个父节点的内容):

final_result = '&'.join(['!'.join(group) for group in unique_groups])

完整可运行示例

把上面的步骤整合起来,用你的示例XML测试:

import xml.etree.ElementTree as ET

# 你的示例XML内容
sample_xml = '''<root> <parent-1> <text>blah-1</text> <properties> <property type="R" id="0005">text-value-A</property> <property type="W" id="0003">text-value-B</property> <property type="H" id="0002">text-value-C</property> <property type="W" id="0008">text-value-D</property> </properties> </parent-1> <parent-2> <text>blah-2</text> <properties> <property type="W" id="0004">text-value-A</property> <property type="H" id="0087">text-value-B</property> </properties> </parent-2> <parent-3> <text>blah-3</text> <properties> <property type="H" id="0087">text-value-C</property> <property type="R" id="0008">text-value-A</property> </properties> </parent-3> <parent-4> <text>blah-4</text> <properties> <property type="H" id="0019">text-value-C</property> <property type="R" id="0060">text-value-A</property> </properties> </parent-4> </root>'''

# 解析XML
root = ET.fromstring(sample_xml)

# 提取父节点文本分组
parent_text_groups = []
for parent in root.findall('.//*[starts-with(local-name(), "parent-")]'):
    prop_texts = [prop.text for prop in parent.find('./properties').findall('property')]
    parent_text_groups.append(prop_texts)

# 去重
unique_groups = []
seen_groups = set()
for group in parent_text_groups:
    group_tuple = tuple(group)
    if group_tuple not in seen_groups:
        seen_groups.add(group_tuple)
        unique_groups.append(group)

# 拼接结果
final_result = '&'.join(['!'.join(group) for group in unique_groups])

print(final_result)
# 输出正好是你要的:text-value-A!text-value-B!text-value-C!text-value-D&text-value-A!text-value-B&text-value-C!text-value-A

为啥你的原思路不行?

你的代码把所有property的文本混进一个列表后去重,不仅丢失了每个父节点的分组信息,还会错误地去掉单个重复的文本(比如如果两个父节点都有text-value-A,原代码会只保留一个),这完全不符合你“剔除整个文本列表重复的父节点”的需求。必须先按父节点分组,再处理去重和拼接,才能达到预期效果。

内容的提问来源于stack exchange,提问作者bt123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:27:20