You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用BeautifulSoup解析XML时如何提取标签内注释文本

问题说明
  • 采集对象为世界银行公开数据集的DDI/XML格式元数据文件
  • 文件中a3_1字段对应的qstnLit标签核心内容被放在XML注释块内,示例内容为:[CDATA[ Please identify which of the following you consider the most important development priorities in Vietnam. (Choose no more than THREE) - Food safety ]]
  • 现有基于BeautifulSoup编写的采集代码仅能提取未被注释的普通文本,无法获取注释块内的目标信息

原有采集代码

from bs4 import BeautifulSoup

file = "VNM_2020_WBCS_v01_M.xml"
import pandas as pd

with open(file, "r", encoding = "utf-8") as f:
                    sauce = f.read()
                    soup = BeautifulSoup(sauce, features = "lxml")

ref = pd.DataFrame()
soup = soup.select("codeBook")[0].select("dataDscr")[0].select("var")

for txt in soup:
                    try:
                        part = pd.DataFrame(data = {"ID"    : [txt.attrs["name"]],
                                                    "Qn"    : [txt.find("qstn").find("qstnlit").text],
                                                    "Lbl"   : [txt.find("labl").text],
                                                    "Max"   : [txt.find(attrs = {"type":"max"}).text],
                                                    "Min"   : [txt.find(attrs = {"type":"min"}).text]
                                                        })
                        ref = ref.append(part)
                    except:
                        pass
实现方法

BeautifulSoup原生支持识别XML中的注释节点,注释内容会被归类为Comment类型,默认调用.text属性时不会提取这类节点的内容,只需在遍历节点时单独判断、提取注释节点内容即可,无需更换解析工具。

  • 从bs4.element导入Comment类型,用于识别注释节点
  • 定位到qstnLit节点后,遍历其下所有子节点:普通文本节点直接提取内容,Comment类型节点同样提取文本,同时清理内容外层多余的CDATA标记
  • 原代码使用裸except会吞掉所有运行错误,建议缩小捕获范围,仅在缺少对应字段时跳过条目;另外pandas新版本已弃用DataFrame.append()方法,替换为pd.concat()运行效率更高、兼容性更好

修改后可正常提取注释内容的代码

from bs4 import BeautifulSoup
from bs4.element import Comment
import pandas as pd

file = "VNM_2020_WBCS_v01_M.xml"

with open(file, "r", encoding = "utf-8") as f:
    sauce = f.read()
    soup = BeautifulSoup(sauce, features = "lxml")

ref = pd.DataFrame()
var_nodes = soup.select("codeBook")[0].select("dataDscr")[0].select("var")

def extract_qstn_text(qstnlit_node):
    content_parts = []
    for child in qstnlit_node.contents:
        if isinstance(child, Comment):
            # 清理注释内容外层的CDATA标识
            clean_text = child.strip().removeprefix("[CDATA[").removesuffix("]]").strip()
            content_parts.append(clean_text)
        else:
            plain_text = child.strip()
            if plain_text:
                content_parts.append(plain_text)
    return " ".join(content_parts)

for node in var_nodes:
    try:
        qstnlit_node = node.find("qstn").find("qstnlit")
        row = pd.DataFrame(data = {
            "ID": [node.attrs["name"]],
            "Qn": [extract_qstn_text(qstnlit_node)],
            "Lbl": [node.find("labl").text],
            "Max": [node.find(attrs={"type":"max"}).text],
            "Min": [node.find(attrs={"type":"min"}).text]
        })
        ref = pd.concat([ref, row], ignore_index=True)
    except AttributeError:
        # 仅跳过缺失必要字段的条目
        pass

内容的提问来源于stack exchange,提问作者Kah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 10:36:43