You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用IMPORTXML提取含productPriceJSON_的DIV及其中discountPrice值

问题解答

1. 正确匹配目标DIV的方法

你之前的XPath匹配失败有两个核心原因:

  • 语法错误:contains() 函数的使用逻辑错误,第一个参数需要指定要匹配的属性(这里是@id),第二个参数是要检索的子串,你之前的写法把判断逻辑放在了@id的赋值位置,结构完全错误
  • 匹配前缀错误:你写的检索字符串是skuPriceJSON_,但实际目标DIV的ID前缀是productPriceJSON_,两者不匹配自然拿不到结果

正确的XPath写法如下,用starts-with匹配前缀更精准,避免匹配到ID中间包含目标字符串的无关元素:

//div[starts-with(@id, "productPriceJSON_")]

如果确实需要包含匹配,用contains写法如下:

//div[contains(@id, "productPriceJSON_")]

2. 提取discountPrice字段的实现方案

目标DIV的内容是标准JSON数组结构,只要把DIV的文本内容取出做JSON解析,就可以直接读取字段值,以下是两种常见场景的实现代码:

浏览器端JavaScript实现

// 匹配所有符合条件的DIV,也可以用上面的XPath执行匹配
const priceDivs = document.querySelectorAll('div[id^="productPriceJSON_"]');
// 遍历提取discountPrice和对应产品ID
const discountPriceList = Array.from(priceDivs).map(div => {
  // 解析JSON内容,注意内容放在数组中,需要取下标0的对象
  const priceInfo = JSON.parse(div.textContent.trim())[0];
  return {
    productId: priceInfo.productId,
    discountPrice: priceInfo.discountPrice
  };
});
// 输出结果
console.log(discountPriceList);

Python爬虫(lxml库)实现

import json
from lxml import etree

# 假设page_source为你爬取到的网页源码
html = etree.HTML(page_source)
# 用XPath匹配所有目标DIV
priceDivs = html.xpath('//div[starts-with(@id, "productPriceJSON_")]')

discountPriceList = []
for div in priceDivs:
    # 取出DIV内的文本内容
    jsonStr = div.xpath('./text()')[0].strip()
    # 解析JSON
    priceInfo = json.loads(jsonStr)[0]
    discountPriceList.append({
        "productId": priceInfo["productId"],
        "discountPrice": priceInfo["discountPrice"]
    })

print(discountPriceList)

内容的提问来源于stack exchange,提问作者b-george

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 12:48:01