Python解析XML链接报错'NoneType' object is not callable及遍历方法求助
问题解决:解析链接报错与数据提取方案
报错原因分析
- 方法误用:
getroot()是xml.etree.ElementTree.ElementTree类的专属方法,你调用的是BeautifulSoup实例的该方法,但BeautifulSoup对象根本没有这个方法,直接触发'NoneType' object is not callable错误。 - 数据类型误解:目标链接返回的不是XML,而是被
<html><body><p>标签包裹的JSON数据,从你打印的response.content和soup输出能看到明显的JSON结构转义字符串。
正确处理步骤
步骤1:提取并解析JSON数据
因为目标链接返回的是JSON,不需要用XML解析工具,正确代码如下:
import requests import json from bs4 import BeautifulSoup url = 'https://www.omicsdi.org/ws/dataset/pride/PXD002885?debug=false' response = requests.get(url) # 提取HTML包裹的JSON字符串 soup = BeautifulSoup(response.content, 'lxml') json_raw = soup.find('p').text # 解析为Python字典 dataset_data = json.loads(json_raw) # 遍历数据(示例:遍历文件版本与文件链接) for file_version in dataset_data.get('file_versions', []): file_groups = file_version.get('files', {}) for file_type, file_url_list in file_groups.items(): print(f"* 文件类型:{file_type}") for url in file_url_list: print(f" - {url}")
步骤2:如果需要处理XML的正确方式
若后续你有真实的XML链接需要解析,使用xml.etree.ElementTree的正确流程如下:
import requests import xml.etree.ElementTree as ET # 替换为真实XML链接 xml_url = "your_valid_xml_link_here" response = requests.get(xml_url) # 解析XML并获取根节点 root = ET.fromstring(response.content) # 遍历entry标签下的所有string标签并打印内容 for entry in root.findall('entry'): for string_tag in entry.findall('string'): if string_tag.text: print(string_tag.text.strip())
内容的提问来源于stack exchange,提问作者Hemant Srivastava
相关产品推荐
相关产品推荐

