You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BS4(lxml)解析XML数据:提取cell_value1无输出问题求助

解决BeautifulSoup提取XML元素属性无输出的问题

你的代码逻辑本身没问题,但没输出的核心原因是XML命名空间导致找不到<Cell>元素,再加上解析器选择不够适配XML场景。我来一步步帮你修复:

问题分析

USDA的这个XML文档带有默认命名空间(根元素会包含类似xmlns="http://www.w3.org/2003/10/Serialization"的属性)。在XML规则里,所有带命名空间的元素,实际标签名是{命名空间URL}Cell,而不是单纯的Cell——这就导致你的soup.find_all("Cell")返回空列表,循环根本没执行,自然没有输出。

另外,处理XML时用专门的xml解析器(或lxml-xml)比通用的lxml解析器更可靠,能更好地识别XML结构。

修复后的代码

方案1:明确指定命名空间查找

from bs4 import BeautifulSoup
import requests

r = requests.get('https://www.usda.gov/oce/commodity/wasde/latest.xml')
data = r.text
# 使用xml解析器处理XML文档
soup = BeautifulSoup(data, "xml")

# 从根元素获取默认命名空间
root_namespace = soup.find().attrs.get('xmlns')
# 查找带命名空间的Cell元素
for cell in soup.find_all(f"{{{root_namespace}}}Cell"):
    # 确保元素存在cell_value1属性再打印,避免潜在报错
    if 'cell_value1' in cell.attrs:
        print(cell['cell_value1'])

方案2:忽略命名空间,匹配标签名后缀

如果不想手动处理命名空间,可以用lambda表达式匹配所有以Cell结尾的标签名(不管前缀):

from bs4 import BeautifulSoup
import requests

r = requests.get('https://www.usda.gov/oce/commodity/wasde/latest.xml')
data = r.text
soup = BeautifulSoup(data, "xml")

# 匹配所有标签名以Cell结尾的元素
for cell in soup.find_all(lambda tag: tag.name.endswith('Cell')):
    if 'cell_value1' in cell.attrs:
        print(cell['cell_value1'])

快速验证步骤

你可以先添加一行代码确认是否找到目标元素:

print(len(soup.find_all(f"{{{root_namespace}}}Cell")))

如果输出大于0,说明元素已被正确定位,后续打印属性就没问题了。

内容的提问来源于stack exchange,提问作者gabe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:25:54