如何在Python/Selenium中获取HTML代码中的SKU编号?
如何在Python/Selenium中获取HTML中的SKU编号
京东搜索页的商品SKU信息通常存储在商品项元素的data-sku属性中,你可以通过以下方式获取:
- 修改元素定位:从定位整个商品列表容器,改为定位列表内的每个商品项(通常是
<li>标签) - 使用
get_attribute()方法提取目标属性值,而非获取元素文本
修改后的完整代码如下:
import time from selenium import webdriver driver = webdriver.Chrome() driver.get('https://search.jd.com/Search?keyword=%E6%9E%9C%E6%B1%81&qrst=1&wq=%E6%9E%9C%E6%B1%81&stock=1&pvid=b86735ca93754d6f96a68a4ee0e187d5&psort=3&click=0') # 执行滚动加载脚本 driver.execute_script(""" (function () { var y = 0; var step = 100; window.scroll(0, 0); function f() { if (y < document.body.scrollHeight) { y += step; window.scroll(0, y); setTimeout(f, 100); } else { window.scroll(0, 0); document.title += "scroll-done"; } } setTimeout(f, 1000); })(); """) print("下拉中...") # 等待滚动完成 while True: if "scroll-done" in driver.title: break else: print("还没有拉到最底端...") time.sleep(3) # 定位每个商品项,并提取data-sku属性值 sku_items = driver.find_elements_by_xpath("//div[@id='J_goodsList']/ul/li") for item in sku_items: sku = item.get_attribute('data-sku') if sku: # 过滤可能为空的情况 print(f"SKU编号:{sku}") driver.quit()
关键说明:
find_elements_by_xpath("//div[@id='J_goodsList']/ul/li"):定位到每个商品的<li>元素,这类元素包含data-sku属性item.get_attribute('data-sku'):直接提取元素的data-sku属性值,这就是HTML中存储的SKU编号- 增加
driver.quit()关闭浏览器,避免资源占用
内容的提问来源于stack exchange,提问作者anderwyang
相关产品推荐
相关产品推荐

