Python/Pandas中如何拼接重复同名XML标签的文本内容
解决方案
你现有代码直接全局遍历所有neighbor节点,没有按所属country父节点分组,因此无法实现同节点下同名标签的文本拼接。调整遍历逻辑,先逐个定位country节点,再收集当前节点下所有neighbor子节点的文本,处理掉标签自带的双引号后用空格拼接即可。
实现代码:
from xml.etree import ElementTree as ET tree = ET.parse('sample.xml') root = tree.getroot() for country in root.iter('country'): neighbor_list = [] for neighbor in country.findall('neighbor'): # 去除文本首尾的空白和自带的双引号 clean_text = neighbor.text.strip().strip('"') neighbor_list.append(clean_text) print(' '.join(neighbor_list))
运行后输出完全匹配预期:
Austria Switzerland Malaysia Costa Rica Colombia
逻辑说明
- 外层循环以
country为遍历单位,天然完成同属节点的分组 - 用
findall('neighbor')仅查找当前country下的子neighbor节点,不会跨节点取值 - 对原始标签内自带的双引号做剥离处理,输出无多余符号
- 用
str.join()拼接字符串,自动适配单邻国、多邻国场景,无需额外写分支判断
内容的提问来源于stack exchange,提问作者OldManSeph
相关产品推荐
相关产品推荐

