Python使用BeautifulSoup解析XML时如何提取标签内注释文本
问题说明
- 采集对象为世界银行公开数据集的DDI/XML格式元数据文件
- 文件中
a3_1字段对应的qstnLit标签核心内容被放在XML注释块内,示例内容为:[CDATA[ Please identify which of the following you consider the most important development priorities in Vietnam. (Choose no more than THREE) - Food safety ]] - 现有基于BeautifulSoup编写的采集代码仅能提取未被注释的普通文本,无法获取注释块内的目标信息
原有采集代码
from bs4 import BeautifulSoup file = "VNM_2020_WBCS_v01_M.xml" import pandas as pd with open(file, "r", encoding = "utf-8") as f: sauce = f.read() soup = BeautifulSoup(sauce, features = "lxml") ref = pd.DataFrame() soup = soup.select("codeBook")[0].select("dataDscr")[0].select("var") for txt in soup: try: part = pd.DataFrame(data = {"ID" : [txt.attrs["name"]], "Qn" : [txt.find("qstn").find("qstnlit").text], "Lbl" : [txt.find("labl").text], "Max" : [txt.find(attrs = {"type":"max"}).text], "Min" : [txt.find(attrs = {"type":"min"}).text] }) ref = ref.append(part) except: pass
实现方法
BeautifulSoup原生支持识别XML中的注释节点,注释内容会被归类为Comment类型,默认调用.text属性时不会提取这类节点的内容,只需在遍历节点时单独判断、提取注释节点内容即可,无需更换解析工具。
- 从
bs4.element导入Comment类型,用于识别注释节点 - 定位到
qstnLit节点后,遍历其下所有子节点:普通文本节点直接提取内容,Comment类型节点同样提取文本,同时清理内容外层多余的CDATA标记 - 原代码使用裸
except会吞掉所有运行错误,建议缩小捕获范围,仅在缺少对应字段时跳过条目;另外pandas新版本已弃用DataFrame.append()方法,替换为pd.concat()运行效率更高、兼容性更好
修改后可正常提取注释内容的代码
from bs4 import BeautifulSoup from bs4.element import Comment import pandas as pd file = "VNM_2020_WBCS_v01_M.xml" with open(file, "r", encoding = "utf-8") as f: sauce = f.read() soup = BeautifulSoup(sauce, features = "lxml") ref = pd.DataFrame() var_nodes = soup.select("codeBook")[0].select("dataDscr")[0].select("var") def extract_qstn_text(qstnlit_node): content_parts = [] for child in qstnlit_node.contents: if isinstance(child, Comment): # 清理注释内容外层的CDATA标识 clean_text = child.strip().removeprefix("[CDATA[").removesuffix("]]").strip() content_parts.append(clean_text) else: plain_text = child.strip() if plain_text: content_parts.append(plain_text) return " ".join(content_parts) for node in var_nodes: try: qstnlit_node = node.find("qstn").find("qstnlit") row = pd.DataFrame(data = { "ID": [node.attrs["name"]], "Qn": [extract_qstn_text(qstnlit_node)], "Lbl": [node.find("labl").text], "Max": [node.find(attrs={"type":"max"}).text], "Min": [node.find(attrs={"type":"min"}).text] }) ref = pd.concat([ref, row], ignore_index=True) except AttributeError: # 仅跳过缺失必要字段的条目 pass
内容的提问来源于stack exchange,提问作者Kah
相关产品推荐
相关产品推荐

