You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从含105条议会发言的XML中提取指定节点字段的技术需求

嘿,我来帮你搞定这份议会发言XML的数据提取需求!针对你那包含105条记录的文件,我准备了两种实用的Python实现方案,都是处理这类结构化XML的主流方法,你可以根据自己的熟悉程度来选:

方案1:用lxml + XPath(精准高效,适合规范的结构化XML)

这是处理XML最常用的组合,XPath能精准定位到你需要的节点,效率很高,尤其适合你这种有固定层级的文件。

首先安装依赖:

pip install lxml

然后是代码示例,已经帮你处理了「无政党字段」的情况,还支持把结果导出成CSV方便后续分析:

from lxml import etree
import csv

# 加载本地XML文件(如果是远程URL,可以用etree.fromstring(requests.get(url).content))
tree = etree.parse("your_parliament_speeches.xml")
root = tree.getroot()

speech_records = []

# 遍历所有发言节点
for spreekbeurt in root.xpath("//spreekbeurt"):
    # 提取发言人姓氏,处理节点不存在的极端情况
    achternaam_nodes = spreekbeurt.xpath(".//spreker/naam/achternaam/text()")
    achternaam = achternaam_nodes[0].strip() if achternaam_nodes else "未知姓氏"
    
    # 提取政党信息,无此字段则标记为「无政党身份」
    partij_nodes = spreekbeurt.xpath(".//spreker/partij/text()")
    partij = partij_nodes[0].strip() if partij_nodes else "无政党身份"
    
    speech_records.append({
        "发言人姓氏": achternaam,
        "政党": partij
    })

# 打印前5条验证结果
print("前5条提取结果:")
for record in speech_records[:5]:
    print(record)

# 导出为CSV文件
with open("parliament_speeches_extracted.csv", "w", newline="", encoding="utf-8") as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=["发言人姓氏", "政党"])
    writer.writeheader()
    writer.writerows(speech_records)
方案2:用BeautifulSoup(灵活容错,适合结构不太规整的XML)

如果你对XPath不太熟悉,BeautifulSoup的API更直观,容错性也更强,就算XML有小瑕疵也能处理。

先安装依赖:

pip install beautifulsoup4 lxml

代码示例如下:

from bs4 import BeautifulSoup
import csv

# 读取本地XML文件
with open("your_parliament_speeches.xml", "r", encoding="utf-8") as xml_file:
    xml_content = xml_file.read()

soup = BeautifulSoup(xml_content, "lxml")

speech_records = []

# 遍历所有发言标签
for spreekbeurt in soup.find_all("spreekbeurt"):
    spreker = spreekbeurt.find("spreker")
    if not spreker:
        continue  # 如果没有发言人节点,跳过这条记录
    
    # 提取姓氏
    naam = spreker.find("naam")
    achternaam = naam.find("achternaam").get_text(strip=True) if naam and naam.find("achternaam") else "未知姓氏"
    
    # 提取政党
    partij = spreker.find("partij").get_text(strip=True) if spreker.find("partij") else "无政党身份"
    
    speech_records.append({
        "发言人姓氏": achternaam,
        "政党": partij
    })

# 验证结果并导出CSV
print("前5条提取结果:")
for record in speech_records[:5]:
    print(record)

with open("parliament_speeches_extracted.csv", "w", newline="", encoding="utf-8") as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=["发言人姓氏", "政党"])
    writer.writeheader()
    writer.writerows(speech_records)

小提示

  • 记得把代码里的your_parliament_speeches.xml替换成你实际的XML文件名
  • 如果XML文件编码不是UTF-8,打开文件时要指定对应的编码(比如encoding='iso-8859-1')
  • 两种方案都支持扩展:如果之后需要提取其他字段(比如发言时间、内容),只需要在循环里添加对应的节点提取逻辑就行

内容的提问来源于stack exchange,提问作者bbilal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:03:53