You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup解析以HTML视图展示的XML文件并提取数据

问题根源

你在浏览器中看到的带有#folder0、div.opened这类元素的页面,是浏览器对XML响应自动渲染生成的交互式可视化结构,并非接口原生返回的内容。你用requests直接请求该接口拿到的是纯XML数据,完全没有对应的HTML节点,因此使用浏览器复制的CSS选择器自然查不到任何结果。
另外你之前直接用XML解析失败,大概率是忽略了该XML内容带官方命名空间,直接按无命名空间的方式查找节点会返回空。

解决方案

首先安装依赖:

pip install lxml

方法1:用BeautifulSoup直接解析XML

代码示例如下:

import requests
from bs4 import BeautifulSoup

url = "https://musicbrainz.org/ws/2/artist/?query=artist:massive-attack"
# 按接口要求添加User-Agent标识,避免被拦截
headers = {
    "User-Agent": "MyDataCrawler/1.0"
}
resp = requests.get(url, headers=headers).text
# 用XML解析器而非HTML解析器
soup = BeautifulSoup(resp, "lxml-xml")

# 查找所有artist节点
artists = soup.find_all("artist")
for artist in artists:
    # 提取对应字段
    artist_name = artist.find("name").get_text() if artist.find("name") else ""
    begin_area = artist.find("begin-area").find("name").get_text() if artist.find("begin-area") and artist.find("begin-area").find("name") else ""
    print(f"歌手名:{artist_name},起源地:{begin_area}")

方法2:用ElementTree解析(需处理命名空间)

代码示例如下:

import requests
import xml.etree.ElementTree as ET

url = "https://musicbrainz.org/ws/2/artist/?query=artist:massive-attack"
headers = {
    "User-Agent": "MyDataCrawler/1.0"
}
resp = requests.get(url, headers=headers).text
root = ET.fromstring(resp)
# 声明XML命名空间
ns = {"mb": "http://musicbrainz.org/ns/mmd-2.0#"}
artists = root.findall(".//mb:artist", ns)
for artist in artists:
    artist_name = artist.find("mb:name", ns).text if artist.find("mb:name", ns) else ""
    begin_area = artist.find("mb:begin-area/mb:name", ns).text if artist.find("mb:begin-area/mb:name", ns) else ""
    print(f"歌手名:{artist_name},起源地:{begin_area}")

你可以先打印resp的内容,就能看到原生返回的XML结构,按照实际结构查找节点即可拿到对应数据。

内容的提问来源于stack exchange,提问作者em77

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 20:24:01