如何用Python从HTML中提取The North Face商品价格?
解决The North Face网站价格提取失败的问题
你遇到的价格提取失败,大概率是因为价格数据是JavaScript动态渲染的——静态HTML源码里并没有直接包含价格元素,而是页面加载后通过JS从后端接口拉取或生成,所以用requests直接获取的HTML里找不到对应的.product-price price元素。以下是两种可行的解决方案:
方案1:从页面初始化脚本中提取价格数据
很多电商网站会在页面的<script>标签里嵌入初始化状态数据(比如window.__INITIAL_STATE__),里面包含完整的商品信息,包括价格。你可以提取这段脚本内容,解析成JSON后获取价格:
import requests import json import re from bs4 import BeautifulSoup as bs # 添加请求头模拟浏览器,降低被识别为爬虫的概率 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } r = requests.get("https://www.thenorthface.it/shop/it/tnf-it/berretto-jim-a5wh?variationId=KS7", headers=headers) contenuto = bs(r.text, "html.parser") # 遍历script标签,找到包含初始化状态的内容 for script in contenuto.find_all("script"): script_content = script.string if script_content and "__INITIAL_STATE__" in script_content: # 用正则匹配出JSON结构 match = re.search(r"window\.__INITIAL_STATE__\s*=\s*({.*?});", script_content, re.DOTALL) if match: state_data = json.loads(match.group(1)) # 从JSON结构中定位价格(若路径失效,可先print(state_data)查看完整结构调整) price = state_data.get("product", {}).get("currentProduct", {}).get("price", {}).get("formatted", "") print(f"价格: {price}") break
方案2:用Selenium等待元素加载完成
之前用Selenium返回空字符串,是因为没有等待元素完全渲染就直接获取。需要通过显式等待确保价格元素加载完成:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() driver.get("https://www.thenorthface.it/shop/it/tnf-it/berretto-jim-a5wh?variationId=KS7") # 最多等待10秒,直到价格元素出现 try: price_element = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "product-price")) ) print(f"价格: {price_element.text.strip()}") finally: driver.quit()
注意事项
- 电商网站的页面结构和初始化数据格式可能变动,若JSON路径失效,可先打印完整的
state_data对象,定位价格对应的具体层级。 - 频繁爬取可能触发反爬机制,建议添加请求间隔,避免IP被封禁。
内容的提问来源于stack exchange,提问作者Pietro Leon
相关产品推荐
相关产品推荐

