如何使用IMPORTXML提取含productPriceJSON_的DIV及其中discountPrice值
问题解答
1. 正确匹配目标DIV的方法
你之前的XPath匹配失败有两个核心原因:
- 语法错误:
contains()函数的使用逻辑错误,第一个参数需要指定要匹配的属性(这里是@id),第二个参数是要检索的子串,你之前的写法把判断逻辑放在了@id的赋值位置,结构完全错误 - 匹配前缀错误:你写的检索字符串是
skuPriceJSON_,但实际目标DIV的ID前缀是productPriceJSON_,两者不匹配自然拿不到结果
正确的XPath写法如下,用starts-with匹配前缀更精准,避免匹配到ID中间包含目标字符串的无关元素:
//div[starts-with(@id, "productPriceJSON_")]
如果确实需要包含匹配,用contains写法如下:
//div[contains(@id, "productPriceJSON_")]
2. 提取discountPrice字段的实现方案
目标DIV的内容是标准JSON数组结构,只要把DIV的文本内容取出做JSON解析,就可以直接读取字段值,以下是两种常见场景的实现代码:
浏览器端JavaScript实现
// 匹配所有符合条件的DIV,也可以用上面的XPath执行匹配 const priceDivs = document.querySelectorAll('div[id^="productPriceJSON_"]'); // 遍历提取discountPrice和对应产品ID const discountPriceList = Array.from(priceDivs).map(div => { // 解析JSON内容,注意内容放在数组中,需要取下标0的对象 const priceInfo = JSON.parse(div.textContent.trim())[0]; return { productId: priceInfo.productId, discountPrice: priceInfo.discountPrice }; }); // 输出结果 console.log(discountPriceList);
Python爬虫(lxml库)实现
import json from lxml import etree # 假设page_source为你爬取到的网页源码 html = etree.HTML(page_source) # 用XPath匹配所有目标DIV priceDivs = html.xpath('//div[starts-with(@id, "productPriceJSON_")]') discountPriceList = [] for div in priceDivs: # 取出DIV内的文本内容 jsonStr = div.xpath('./text()')[0].strip() # 解析JSON priceInfo = json.loads(jsonStr)[0] discountPriceList.append({ "productId": priceInfo["productId"], "discountPrice": priceInfo["discountPrice"] }) print(discountPriceList)
内容的提问来源于stack exchange,提问作者b-george
相关产品推荐
相关产品推荐

