如何用Python获取沃尔玛网页完整HTML以查验PS5商品库存状态
问题原因梳理
urllib获取内容不全的原因
沃尔玛页面采用客户端动态渲染逻辑,且配置了反爬机制,urllib发起的原生请求未携带合法浏览器UA、Cookie等校验参数,会被反爬系统拦截,返回的是仅包含基础样式的拦截页面,自然无法检索到库存相关的目标元素。Selenium报错原因
- 两个报错本质为同一问题:本地未安装匹配Chrome浏览器版本的ChromeDriver驱动,或是驱动文件未加入系统PATH环境变量,Selenium无法定位到驱动程序。
- 代码本身存在语法逻辑错误:
find_element_by_tag_name仅会返回单个匹配的link标签,你要获取所有link标签需要用find_elements_by_tag_name- link标签的属性值不存在于text字段中,你要获取href属性需要调用
get_attribute('href')方法,直接读取i.text只能拿到空值。
修复方案
步骤1:修复Selenium运行环境
- 打开本地Chrome浏览器,在「设置-关于Chrome」中查看当前浏览器版本
- 下载与浏览器版本完全匹配的ChromeDriver驱动文件,将驱动文件放到Python安装根目录,或者在代码中直接指定驱动文件的绝对路径,无需修改系统PATH。
步骤2:使用修正后的代码实现库存检测
这里直接提供可运行的完整代码示例:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By import time # 此处替换为你本地chromedriver.exe的实际路径,比如r'C:\Users\xxx\Downloads\chromedriver.exe' driver_path = r'你的chromedriver实际路径' s = Service(driver_path) options = webdriver.ChromeOptions() options.add_argument('--headless') options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') # 替换为你当前浏览器对应的UA,避免被反爬识别 options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/116.0.0.0 Safari/537.36') browser = webdriver.Chrome(service=s, options=options) browser.set_page_load_timeout(30) try: browser.get('https://www.walmart.com/ip/Sony-PlayStation-5-PS5-Digital-Edition/607045762') # 等待页面渲染完成,可根据网络情况调整等待时长 time.sleep(5) # 直接检索页面源码匹配库存标识,判断逻辑和你需求一致 page_source = browser.page_source if '//schema.org/InStock' in page_source: print("PS5数字版有货") elif 'Free delivery Arrives by' in page_source: print("PS5数字版有货") else: print("PS5数字版无货") finally: browser.quit()
补充说明
如果需要持续检测,可在代码外层加循环,设置固定间隔时间发起请求即可,注意不要请求过于频繁触发反爬封禁,建议间隔10分钟以上查询一次。
内容的提问来源于stack exchange,提问作者cerealfish
相关产品推荐
相关产品推荐

