You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup转换XML至JSON时H3与Standard标签配对异常问题

问题根因

原代码使用zip(header, sub)对两类标签的查询结果做按下标一一配对,完全忽略了XML文档的节点顺序关系:单个H3标签后跟随多个同级Standard标签,这种硬配对逻辑只会给每个H3匹配到1个Standard标签,剩余的Standard内容会全部丢失,最终结果不符合预期。

实现逻辑

不再单独提取两类标签后错位配对,改为按文档顺序遍历节点:

  • 遇到H3标签时,先将上一个H3及其收集到的所有Standard内容存入结果,再初始化新的H3分组
  • 遇到Standard标签时,直接将内容追加到当前激活的H3分组的内容池中
  • 遍历结束后补充存入最后一个H3分组的内容
实现代码
import re

# 保留原代码的正则规则
extract2 = re.compile(r"[A-Z][a-z]\w*")
result = []
current_h3 = None
current_standard_text = ""

# 遍历所有同级节点,跳过无意义的空文本节点
for node in bs_content.find_all(recursive=False):
    if node.name == "h3":
        # 上一个H3的内容收集完成,存入结果
        if current_h3 is not None:
            req_id = current_h3.text.strip()
            result.append({req_id: current_standard_text})
        # 初始化新H3的收集容器
        current_h3 = node
        current_standard_text = ""
    elif node.name == "standard":
        # 把当前Standard内容追加到所属H3的内容下
        if current_h3 is not None:
            current_standard_text += node.text

# 补存最后一个H3的内容
if current_h3 is not None:
    req_id = current_h3.text.strip()
    result.append({req_id: current_standard_text})
自定义调整说明
  • 如果需要提取H3文本里的编号作为字典键,把req_id的赋值语句替换为原逻辑req_id = str.strip(re.split(extract2, current_h3.text)[0])即可
  • 如果拼接Standard内容时需要加空格、换行等分隔符,修改追加内容的逻辑即可,例如加空格分隔:current_standard_text += " " + node.text
  • 如果H3和Standard标签嵌套在其他父标签下,将bs_content.find_all(recursive=False)替换为对应父节点的find_all(recursive=False)调用即可,核心遍历逻辑不变
输出效果

最终生成的result结构如下,完全匹配预期格式:

[
    {"ISSS.1.A1Acceptance of Overall(B)": "An organisation's Top Management.The Top Management MUST define."},
    {"ISS.2.A2Acceptance of Overall(C)": "An organisation's Top.Top Management."},
    {"ISS.2.2Acceptance of Overall(D)": "An organisation's Top resource.Top Management resource."}
]

内容的提问来源于stack exchange,提问作者sagargahalod

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 23:18:19