You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup迭代获取属性值仅返回首个元素的问题排查

问题:无法获取XML中所有图片路径

我通过以下代码读取本地HTML/XML文件爬取图片路径:

from bs4 import BeautifulSoup as bs

path_xml = r"..."

content = []

with open(path_xml, "r") as file:
    content = file.readlines()

content = "".join(content)
bs_content = bs(content, "html.parser")

bilder = bs_content.find_all("bilder")

def get_str_bild(match):
    test = match.findChildren("b")

    for x in range(len(test)): # 此处存在问题(无法获取test中的所有元素)
 
        return test[x].get("d")

for b in bilder:
    if b.b: 
        print(get_str_bild(b))

当前输出仅每个<bilder>块的首个图片路径:

L3357U00_002120.jpg
L3357U00_002140.jpg
L3357U00_002160.jpg

XML文件结构(每个<Bilder>节点下包含多个<B>子元素):

<Bilder>
    <B Nr="1" D="L3357U00_002120.jpg"/>
    <B Nr="2" D="L3357U00_002120.jpg"/>
    <B Nr="3" D="L3357U00_002120.jpg"/>
    <B Nr="4" D="L3357U00_002120.jpg"/>
    <B Nr="9" D="L3357U00_002120.jpg"/>
    <B Nr="1" D="L3357U00_002130.jpg"/>
    <B Nr="2" D="L3357U00_002130.jpg"/>
    <B Nr="3" D="L3357U00_002130.jpg"/>
    <B Nr="4" D="L3357U00_002130.jpg"/>
    <B Nr="9" D="L3357U00_002130.jpg"/>
</Bilder>

错误原因

问题出在get_str_bild函数的for循环:

  • 循环第一次迭代就执行return语句,函数直接终止,只会返回第一个<B>元素的D属性值,后续元素根本没机会处理。

修正方案

方案1:修改函数返回所有图片路径列表

调整函数逻辑,收集所有<B>元素的D属性后再返回,最后遍历打印:

from bs4 import BeautifulSoup as bs

path_xml = r"..."

# 直接用read()替代readlines()+join,更简洁
with open(path_xml, "r") as file:
    content = file.read()

bs_content = bs(content, "html.parser")

bilder = bs_content.find_all("bilder")

def get_str_bild(match):
    test = match.findChildren("b")
    # 用列表推导式收集所有D属性值
    return [item.get("d") for item in test]

for b in bilder:
    if b.b: 
        # 遍历列表打印每个路径
        for img_path in get_str_bild(b):
            print(img_path)

方案2:简化代码,移除多余函数

直接在主循环中处理每个<B>元素,省略额外函数:

from bs4 import BeautifulSoup as bs

path_xml = r"..."

with open(path_xml, "r") as file:
    content = file.read()

bs_content = bs(content, "html.parser")

# 按<Bilder>分组处理所有子元素
for bilder_block in bs_content.find_all("bilder"):
    for b_tag in bilder_block.find_all("b"):
        print(b_tag.get("d"))

额外提示

XML标签本身大小写敏感,但你使用的html.parser是大小写不敏感的,所以find_all("bilder")能匹配<Bilder>、find_all("b")能匹配<B>。如果后续改用XML专用解析器(如lxml-xml),需要写成完全匹配的标签名("Bilder"和"B")。

内容的提问来源于stack exchange,提问作者maggTech

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 19:10:18