使用Beautiful Soup与Python爬取页面时无法获取指定价格元素
解决无法抓取1mg产品价格的问题
问题原因分析
- 动态内容渲染:该页面的价格元素是通过JavaScript动态加载的,使用
requests获取的静态HTML中并不包含这个价格节点,因此BeautifulSoup无法定位到它。 - 动态生成的class名:你使用的完整class名(
PriceBoxPlanOption__offer-price___3v9x8 PriceBoxPlanOption__offer-price-cp___2QPU_)包含随机后缀,这类class通常是前端构建工具(如Webpack)生成的,页面更新后后缀可能变化,导致精确匹配失效。
解决方案
方案1:优化静态HTML的选择器(仅当价格存在于静态HTML时有效)
如果静态HTML中实际存在价格元素,只是class名易变,可以使用部分class匹配来定位元素:
import requests from bs4 import BeautifulSoup url = "https://www.1mg.com/otc/iodex-ultra-gel-otc716295" response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') try: # 匹配包含指定前缀的class,忽略随机后缀 price_span = soup.find("span", class_=lambda c: c and "PriceBoxPlanOption__offer-price" in c) if price_span: price = price_span.string.strip().replace(',', '').replace('₹', '') else: price = "NA" except AttributeError: price = "NA" print("Products price = ", price)
方案2:使用浏览器渲染工具(解决动态加载问题)
如果价格是JavaScript动态渲染的,需要用Selenium模拟浏览器加载页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 初始化Chrome浏览器(需提前安装ChromeDriver并配置环境变量) driver = webdriver.Chrome() driver.get("https://www.1mg.com/otc/iodex-ultra-gel-otc716295") try: # 等待价格元素加载完成,最多等待10秒 price_element = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "span[class*='PriceBoxPlanOption__offer-price']")) ) price = price_element.text.strip().replace(',', '').replace('₹', '') finally: # 关闭浏览器 driver.quit() print("Products price = ", price)
额外提示
- 可以先通过浏览器的「查看页面源代码」(Ctrl+U)搜索价格相关文本,确认静态HTML中是否存在该内容,以此判断是否需要使用浏览器渲染工具。
- 部分网站会对爬虫进行反爬限制,使用Selenium时建议添加请求头、设置随机延迟,避免被封禁。
内容的提问来源于stack exchange,提问作者Vipin Kumar
相关产品推荐
相关产品推荐

