You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的Selenium爬取Jiomart 1kg洋葱价格遇报错求助

解决JioMart价格爬取的Selenium报错问题

错误1:headless属性弃用警告

  • 原因:Selenium 4.8+版本后,options.headless = True的写法被官方弃用,需改用命令行参数配置无头模式。
  • 解决方法:替换为options.add_argument('--headless=new')(适配新版Chrome的推荐写法),或兼容旧版的options.add_argument('--headless')。

错误2:'WebDriver'对象无'find_element_by_xpath'属性

  • 原因:Selenium 4.x版本移除了find_element_by_xpath、find_element_by_id等旧版定位方法,统一使用find_element()配合By类指定定位策略。
  • 解决方法:
    1. 导入By类:from selenium.webdriver.common.by import By
    2. 将所有find_element_by_*方法替换为find_element(By.XXX, 定位表达式),例如driver.find_element(By.XPATH, "//input[@id='areaPinCode']")

额外优化:用显式等待替代固定sleep

time.sleep是固定时长等待,易因页面加载延迟导致元素查找失败。推荐使用WebDriverWait实现显式等待,等待元素可交互后再操作,提升代码稳定性。

修正后的完整Selenium代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

# 配置浏览器选项
options = Options()
# 启用新版无头模式
options.add_argument('--headless=new')
# 添加用户代理,避免被识别为爬虫
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")

# 启动Chrome浏览器
driver_path = "C:/chromedriver/chromedriver"
service = Service(executable_path=driver_path)
driver = webdriver.Chrome(service=service, options=options)

try:
    # 访问JioMart首页
    driver.get("https://www.jiomart.com/")

    # 显式等待邮编输入框加载完成并输入有效邮编
    pincode_box = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.XPATH, "//input[@id='areaPinCode']"))
    )
    pincode_box.clear()
    pincode_box.send_keys("110006")
    pincode_box.submit()
    time.sleep(2)  # 等待邮编验证页面跳转

    # 显式等待搜索框加载完成并搜索1kg洋葱
    search_bar = WebDriverWait(driver, 10).until(
        EC.element_to_be_clickable((By.XPATH, "//input[@id='searchbar']"))
    )
    search_bar.clear()
    search_bar.send_keys("onion 1 kg")
    search_bar.submit()
    time.sleep(2)  # 等待搜索结果加载

    # 查找目标产品并提取价格
    product = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.XPATH, "//div[@data-name='Onion 1 Kg']"))
    )
    price = product.find_element(By.CLASS_NAME, "sp__dp").text
    print(f"The price of Onion 1 Kg is {price}")

except Exception as e:
    print(f"Error: {str(e)}")
    print("The product was not found on JioMart")

finally:
    # 确保浏览器关闭
    driver.quit()

关于BeautifulSoup方案的补充

若requests+BeautifulSoup方案无法找到产品,可能是JioMart页面动态加载或URL参数变更导致,可尝试:

  • 验证当前JioMart洋葱分类的实际URL是否正确
  • 使用requests.Session()维持会话,模拟浏览器Cookie状态
  • 补充更多请求头(如Accept-Language、Referer)模拟真实用户请求

内容的提问来源于stack exchange,提问作者Vishwajeet Chandor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 20:15:26