使用BeautifulSoup转换XML至JSON时H3与Standard标签配对异常问题
问题根因
原代码使用zip(header, sub)对两类标签的查询结果做按下标一一配对,完全忽略了XML文档的节点顺序关系:单个H3标签后跟随多个同级Standard标签,这种硬配对逻辑只会给每个H3匹配到1个Standard标签,剩余的Standard内容会全部丢失,最终结果不符合预期。
实现逻辑
不再单独提取两类标签后错位配对,改为按文档顺序遍历节点:
- 遇到H3标签时,先将上一个H3及其收集到的所有Standard内容存入结果,再初始化新的H3分组
- 遇到Standard标签时,直接将内容追加到当前激活的H3分组的内容池中
- 遍历结束后补充存入最后一个H3分组的内容
实现代码
import re # 保留原代码的正则规则 extract2 = re.compile(r"[A-Z][a-z]\w*") result = [] current_h3 = None current_standard_text = "" # 遍历所有同级节点,跳过无意义的空文本节点 for node in bs_content.find_all(recursive=False): if node.name == "h3": # 上一个H3的内容收集完成,存入结果 if current_h3 is not None: req_id = current_h3.text.strip() result.append({req_id: current_standard_text}) # 初始化新H3的收集容器 current_h3 = node current_standard_text = "" elif node.name == "standard": # 把当前Standard内容追加到所属H3的内容下 if current_h3 is not None: current_standard_text += node.text # 补存最后一个H3的内容 if current_h3 is not None: req_id = current_h3.text.strip() result.append({req_id: current_standard_text})
自定义调整说明
- 如果需要提取H3文本里的编号作为字典键,把
req_id的赋值语句替换为原逻辑req_id = str.strip(re.split(extract2, current_h3.text)[0])即可 - 如果拼接Standard内容时需要加空格、换行等分隔符,修改追加内容的逻辑即可,例如加空格分隔:
current_standard_text += " " + node.text - 如果H3和Standard标签嵌套在其他父标签下,将
bs_content.find_all(recursive=False)替换为对应父节点的find_all(recursive=False)调用即可,核心遍历逻辑不变
输出效果
最终生成的result结构如下,完全匹配预期格式:
[ {"ISSS.1.A1Acceptance of Overall(B)": "An organisation's Top Management.The Top Management MUST define."}, {"ISS.2.A2Acceptance of Overall(C)": "An organisation's Top.Top Management."}, {"ISS.2.2Acceptance of Overall(D)": "An organisation's Top resource.Top Management resource."} ]
内容的提问来源于stack exchange,提问作者sagargahalod
相关产品推荐
相关产品推荐

