You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用HTML Agility Pack在单循环中解析电商商品的名称与价格

没问题!要在一次循环里同时抓取商品名称和价格,核心思路是先定位到所有的商品块div,然后在每个商品块内部分别提取名称和价格——这样既不用分开遍历两次,效率更高,还能避免出现名称和价格错位的问题(比如某个商品缺失字段时,不会把下一个商品的价格匹配到当前商品)。

下面我用两种常用的解析工具给你写示例:

Python + BeautifulSoup 示例

首先假设你的网页结构是这样的(电商网站的商品目录结构通常类似):

<div class="product-catalog">
  <div class="product-item">
    <h3 class="product-name">无线蓝牙耳机</h3>
    <span class="product-price">¥199</span>
  </div>
  <div class="product-item">
    <h3 class="product-name">智能手表</h3>
    <span class="product-price">¥399</span>
  </div>
  <!-- 可能存在缺失字段的商品 -->
  <div class="product-item">
    <h3 class="product-name">便携充电宝</h3>
  </div>
</div>

对应的解析代码:

from bs4 import BeautifulSoup

# 这里替换成你实际获取到的网页HTML内容
html_content = """
<div class="product-catalog">
  <div class="product-item">
    <h3 class="product-name">无线蓝牙耳机</h3>
    <span class="product-price">¥199</span>
  </div>
  <div class="product-item">
    <h3 class="product-name">智能手表</h3>
    <span class="product-price">¥399</span>
  </div>
  <div class="product-item">
    <h3 class="product-name">便携充电宝</h3>
  </div>
</div>
"""

# 初始化解析器
soup = BeautifulSoup(html_content, 'html.parser')

# 第一步:定位到商品目录容器
catalog_container = soup.find('div', class_='product-catalog')
# 第二步:获取所有独立的商品块
all_product_items = catalog_container.find_all('div', class_='product-item')

# 第三步:一次循环处理每个商品块
for item in all_product_items:
    # 在当前商品块内提取名称(处理可能的空值)
    product_name = item.find('h3', class_='product-name').get_text(strip=True) if item.find('h3', class_='product-name') else "名称未知"
    # 在当前商品块内提取价格(处理可能的空值)
    product_price = item.find('span', class_='product-price').get_text(strip=True) if item.find('span', class_='product-price') else "价格未知"
    
    # 输出或存储结果
    print(f"商品:{product_name} | 价格:{product_price}")

关键细节:

  • 用item.find()而不是soup.find():确保只在当前商品块的范围内查找元素,不会跨商品匹配
  • 空值判断:避免某些商品缺失名称/价格时抛出异常
  • 一次循环完成两个字段的提取,逻辑更紧凑

JavaScript + Cheerio 示例(适合Node.js环境)

如果是用Node.js做爬虫,Cheerio是常用的HTML解析库,实现思路完全一致:

const cheerio = require('cheerio');

// 替换为实际网页HTML
const htmlContent = `
<div class="product-catalog">
  <div class="product-item">
    <h3 class="product-name">无线蓝牙耳机</h3>
    <span class="product-price">¥199</span>
  </div>
  <div class="product-item">
    <h3 class="product-name">智能手表</h3>
    <span class="product-price">¥399</span>
  </div>
</div>
`;

// 加载HTML
const $ = cheerio.load(htmlContent);

// 遍历每个商品块,同时提取名称和价格
$('.product-catalog .product-item').each((_, element) => {
  const name = $(element).find('.product-name').text().trim() || '名称未知';
  const price = $(element).find('.product-price').text().trim() || '价格未知';
  
  console.log(`商品:${name} | 价格:${price}`);
});

核心总结

不管用哪种工具,核心逻辑都是:

  1. 先把所有商品块一次性抓取出来
  2. 对每个商品块,把它作为独立的上下文,在内部提取需要的所有字段
  3. 处理可能的异常情况(比如字段缺失)

这样就能在一次循环里完成名称和价格的解析,既高效又安全。

内容的提问来源于stack exchange,提问作者AntohaY

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:17:06