You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup遍历父DIV下多同class span仅返回首个记录的问题

问题分析与解决方案

问题根源

你的代码逻辑存在两个关键问题:

  1. soup.find_all(class_="product_constant_fields") 只会返回一个元素(页面中仅存在一个class为product_constant_fields的父div),外层循环仅执行一次。
  2. 循环内部使用item.find(class_="field_name")和item.find(class_="field_value"),find()方法仅返回第一个匹配的标签,因此只能拿到第一组字段。

修正后的代码

from bs4 import BeautifulSoup

# 仅需创建一次BeautifulSoup实例,选择一个解析器即可
soup = BeautifulSoup(response.text, 'lxml')
results = []

# 定位到唯一的product_constant_fields容器
product_container = soup.find(class_="product_constant_fields")
# 获取容器内所有的width-auto子div,每个div对应一组字段
field_groups = product_container.find_all(class_="width-auto")

for group in field_groups:
    parsed = {}
    # 在当前div内查找对应的字段名和值
    field_name = group.find(class_="field_name")
    parsed["field_name"] = field_name.text.strip()  # 去除文本前后空白字符
    field_value = group.find(class_="field_value")
    parsed["field_value"] = field_value.text.strip()
    results.append(parsed)

print(results)

代码说明

  • 移除了重复创建BeautifulSoup的冗余代码,解析器三选一即可(lxml性能最优,html.parser无需额外安装)。
  • 先定位父容器,再遍历容器内的每个width-auto子div,确保每组字段都被处理。
  • 使用strip()去除文本前后的空白字符,让结果更整洁(可根据实际需求调整是否保留)。

预期输出

运行后将得到你期望的完整字段列表:

[
{'field_name': 'Item #:', 'field_value': 'AB11223344'},
{'field_name': 'Brand:', 'field_value': 'Johns'},
{'field_name': 'UPC#:', 'field_value': '12345678901234'},
{'field_name': 'UNSPSC:', 'field_value': '12345678'},
{'field_name': 'ManufacturerNo:', 'field_value': '1234567'},
{'field_name': 'Alternate MFG #:', 'field_value': '87654321'}
]

内容的提问来源于stack exchange,提问作者Kumanan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 11:15:06