You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于单个<zs:record>解析MARCXML数据而非整文件提取?

解决方案

用BeautifulSoup实现

核心思路是遍历每个<zs:record>节点,在单条记录范围内提取目标字段,而非全局查找所有datafield,天然保证同一条记录的字段关联。

代码示例

from bs4 import BeautifulSoup
import requests

# 获取SRU接口数据(也可读取本地XML文件)
url = "https://sru.k10plus.de/opac-de-627!rec=1?version=1.1&operation=searchRetrieve&query=pica.tit%3DGeschichten%20aus%20unserer%20Zeit+and+pica.all%3DHotz,%20Karl&maximumRecords=100&recordSchema=marcxml"
response = requests.get(url)
soup = BeautifulSoup(response.content, features="xml")

# 定义要提取的目标字段
target_tags = ["245", "250", "583", "924"]

all_records = []
# 遍历每个zs:record节点(XML带命名空间,需用完整标签名查找)
for record in soup.find_all("{http://www.loc.gov/zing/srw/}record"):
    # 初始化当前记录的字段容器,支持多字段重复
    record_data = {tag: [] for tag in target_tags}
    
    # 在当前记录内查找所有datafield
    datafields = record.find_all("datafield")
    for df in datafields:
        tag = df.get("tag")
        if tag in target_tags:
            # 提取子字段内容(可按需调整,比如单独提取$a/$b等子字段)
            subfield_content = " ".join([sf.get_text() for sf in df.find_all("subfield")])
            record_data[tag].append(subfield_content)
    
    all_records.append(record_data)

# 输出结果示例
for idx, rec in enumerate(all_records):
    print(f"记录 {idx+1}:")
    for tag, values in rec.items():
        print(f"  {tag}: {values if values else '无此字段'}")

用lxml etree实现

etree处理XML效率更高,尤其适合大体积数据,核心逻辑同样是按记录节点遍历,用XPath精准定位字段:

代码示例

from lxml import etree
import requests

url = "https://sru.k10plus.de/opac-de-627!rec=1?version=1.1&operation=searchRetrieve&query=pica.tit%3DGeschichten%20aus%20unserer%20Zeit+and+pica.all%3DHotz,%20Karl&maximumRecords=100&recordSchema=marcxml"
response = requests.get(url)
root = etree.fromstring(response.content)

# 定义XML命名空间,对应前缀
ns = {
    "zs": "http://www.loc.gov/zing/srw/",
    "marc": "http://www.loc.gov/MARC21/slim"
}

target_tags = ["245", "250", "583", "924"]
all_records = []

# 用XPath遍历所有zs:record节点
for record in root.xpath("//zs:record", namespaces=ns):
    record_data = {tag: [] for tag in target_tags}
    
    # 遍历每个目标字段,在当前记录内查找
    for tag in target_tags:
        datafields = record.xpath(f".//marc:datafield[@tag='{tag}']", namespaces=ns)
        for df in datafields:
            # 提取子字段文本内容
            subfield_text = " ".join(df.xpath(".//marc:subfield/text()", namespaces=ns))
            record_data[tag].append(subfield_text)
    
    all_records.append(record_data)

# 输出结果示例
for idx, rec in enumerate(all_records):
    print(f"记录 {idx+1}:")
    for tag, values in rec.items():
        print(f"  {tag}: {values if values else '无此字段'}")

内容的提问来源于stack exchange,提问作者WorldTeacher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 07:45:24