元素存在但Beautiful Soup返回None:Officeworks AirPods价格爬取失败
问题:无法提取Officeworks网站AirPods的价格
尝试获取Officeworks网站上苹果AirPods的价格,但代码执行后返回None,尽管目标元素在页面中存在。使用的代码如下:
import requests from bs4 import BeautifulSoup def PriceOfficeWorks(URL): HEADERS = { 'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/66.0.3359.181 Safari/537.36', 'Accept-Language': 'en-US, en;q=0.5' } page = requests.get(URL, headers=HEADERS, timeout=10) soup = BeautifulSoup(page.content, "html.parser") price = soup.find('div', class_='sc-TOsTZ jxBJfl') finalprice = price.find('span') return finalprice print(PriceOfficeWorks("https://www.officeworks.com.au/shop/officeworks/p/apple-airpods-with-charging-case-2nd-gen-amv7n2zaa"))
原因分析
- 动态内容渲染:Officeworks的商品价格可能通过JavaScript动态加载,直接用
requests获取的静态HTML中不包含目标元素的真实内容。 - 动态类名不可靠:代码中使用的
sc-TOsTZ jxBJfl是前端框架生成的动态类名,这类名称会随页面更新或渲染逻辑变化,无法作为稳定的定位依据。
解决方案
方法一:利用静态HTML中的稳定属性定位
检查页面静态源码,可通过data-testid这类稳定的标识定位价格元素,Officeworks的商品价格通常带有data-testid="product-price"属性:
import requests from bs4 import BeautifulSoup def PriceOfficeWorks(URL): HEADERS = { 'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept-Language': 'en-US, en;q=0.5' } page = requests.get(URL, headers=HEADERS, timeout=10) soup = BeautifulSoup(page.content, "html.parser") # 使用稳定的data-testid属性定位价格 price_element = soup.find('span', attrs={'data-testid': 'product-price'}) return price_element.text.strip() if price_element else None print(PriceOfficeWorks("https://www.officeworks.com.au/shop/officeworks/p/apple-airpods-with-charging-case-2nd-gen-amv7n2zaa"))
方法二:用Selenium处理动态加载内容
如果静态HTML中确实没有价格数据,可通过Selenium模拟浏览器加载完整页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options import time def PriceOfficeWorks(URL): chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--user-agent=Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36') driver = webdriver.Chrome(options=chrome_options) driver.get(URL) time.sleep(2) # 等待页面动态内容加载完成 try: price_element = driver.find_element(By.CSS_SELECTOR, '[data-testid="product-price"]') return price_element.text.strip() finally: driver.quit() print(PriceOfficeWorks("https://www.officeworks.com.au/shop/officeworks/p/apple-airpods-with-charging-case-2nd-gen-amv7n2zaa"))
内容的提问来源于stack exchange,提问作者lopez__elian223
相关产品推荐
相关产品推荐

