You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup抓取网页中Product Dimensions字段信息

用BeautifulSoup提取Product Dimensions字段实现方案

以下是针对电商类页面规格字段的通用抓取代码,适配绝大多数常见的页面结构:

核心实现逻辑

from bs4 import BeautifulSoup
import re

# 此处soup为你已初始化完成的BeautifulSoup对象
product_info = {}

# 适配列表/自由布局类结构
for text_node in soup.find_all(string=True):
    key = text_node.strip()
    if key == "Product Dimensions":
        # 提取相邻下一个有效节点的文本,清洗多余空格换行
        dim_value = text_node.find_next().get_text(strip=True)
        dim_value = re.sub(r'\s+', ' ', dim_value)
        product_info["Product Dimensions"] = dim_value
        break

# 若上面逻辑未匹配到,补充适配表格结构
if not product_info:
    for tr in soup.find_all("tr"):
        cells = tr.find_all("td")
        if len(cells) < 2:
            continue
        key = cells[0].get_text(strip=True)
        if key == "Product Dimensions":
            dim_value = cells[1].get_text(strip=True)
            dim_value = re.sub(r'\s+', ' ', dim_value)
            product_info["Product Dimensions"] = dim_value
            break

print(product_info)

输出结果

运行后即可得到你需要的结构化数据格式:
{'Product Dimensions': '102.87 x 40.64 x 11.43 cm; 16.33 Kilograms'}

  • 如果你的页面有特殊的标签嵌套规则,可以调整find_next()的参数,指定具体标签名(比如find_next("span"))提升匹配精度

内容的提问来源于stack exchange,提问作者D_S

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 03:06:02