You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析XML链接报错'NoneType' object is not callable及遍历方法求助

问题解决:解析链接报错与数据提取方案

报错原因分析

  1. 方法误用:getroot()是xml.etree.ElementTree.ElementTree类的专属方法,你调用的是BeautifulSoup实例的该方法,但BeautifulSoup对象根本没有这个方法,直接触发'NoneType' object is not callable错误。
  2. 数据类型误解:目标链接返回的不是XML,而是被<html><body><p>标签包裹的JSON数据,从你打印的response.content和soup输出能看到明显的JSON结构转义字符串。

正确处理步骤

步骤1:提取并解析JSON数据

因为目标链接返回的是JSON,不需要用XML解析工具,正确代码如下:

import requests
import json
from bs4 import BeautifulSoup

url = 'https://www.omicsdi.org/ws/dataset/pride/PXD002885?debug=false'
response = requests.get(url)

# 提取HTML包裹的JSON字符串
soup = BeautifulSoup(response.content, 'lxml')
json_raw = soup.find('p').text

# 解析为Python字典
dataset_data = json.loads(json_raw)

# 遍历数据(示例:遍历文件版本与文件链接)
for file_version in dataset_data.get('file_versions', []):
    file_groups = file_version.get('files', {})
    for file_type, file_url_list in file_groups.items():
        print(f"* 文件类型:{file_type}")
        for url in file_url_list:
            print(f"  - {url}")

步骤2:如果需要处理XML的正确方式

若后续你有真实的XML链接需要解析,使用xml.etree.ElementTree的正确流程如下:

import requests
import xml.etree.ElementTree as ET

# 替换为真实XML链接
xml_url = "your_valid_xml_link_here"
response = requests.get(xml_url)

# 解析XML并获取根节点
root = ET.fromstring(response.content)

# 遍历entry标签下的所有string标签并打印内容
for entry in root.findall('entry'):
    for string_tag in entry.findall('string'):
        if string_tag.text:
            print(string_tag.text.strip())

内容的提问来源于stack exchange,提问作者Hemant Srivastava

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 02:54:22