使用Python+Beautiful Soup爬取Newegg GPU价格遇HTML解析问题
解决Newegg网页爬取时HTML结构缺失的问题
问题背景
用Beautiful Soup (BS4)爬取Newegg的GPU价格,请求URL能拿到响应,但解析后的文档只有脚本标签,没有产品价格相关的预期HTML结构。
原代码
from bs4 import BeautifulSoup import requests url = "https://www.newegg.com/evga-geforce-rtx-3080-ti-12g-p5-3967-kr/p/N82E16814487547" result = requests.get(url) print(result.text) doc = BeautifulSoup(result.text, "html.parser") print(doc.prettify())
异常输出
</script> <script defer=""> window.__neweggState__ = {"country":{"name":"United States","alpha2":"us","alpha3":"USA","geoLocation":"North America","currency":"USD"},"user":{"nvtc":"string","contactWith":"string","loginName":"string","accessToken":"string","lastVisitTime":1,"loginId":"string","isEggExpert":true,"loginToken":"string","isPremier":true,"isLogin":true},"domains":{"WWW":"www.newegg.com","SSL":"secure.newegg.com","MOBILESSL":"secure.m.newegg.com","COM":"www.newegg.com","SecureCOM":"secure.newegg.com","CA":"www.newegg.ca","Dynamic":"newegg.com"},"currency":{"extraUnit":"","currencyCode":"USD","countryCode":"USA","unit":"$","decimalDigits":2,"groupSizes":3,"decimalSeparator":".","groupSeparator":",","positivePattern":0,"supportDecimal":true}} </script> <script defer=""> window.__pageInfo__ = {"params":{"keyword":"evga-geforce-rtx-3080-ti-12g-p5-3967-kr","parentItem":"N82E16814487547"},"query":{},"routeName":"Product","isWWWDomain":true,"imageName":"product","hostname":"www.newegg.com","theme":null} </script> <script defer=""> window.__langResouce__ = {} </script> <iframe id="cross_storage_www" src="https://www.newegg.com/api/storageHub" style="display:none"> </iframe> <iframe id="cross_storage_ssl" src="https://secure.newegg.com/api/storageHub" style="display:none"> </iframe> <script defer="" src="https://c1.neweggimages.com/WebResource/Scripts/WWW/vendor~product~ProductDetail-d142e22f.js"> </script> <script defer="" src="https://c1.neweggimages.com/WebResource/Scripts/WWW/common~product~ProductDetail-8045cd77.js"> </script> <script defer="" src="https://c1.neweggimages.com/WebResource/Scripts/WWW/ProductDetail-49788263.js"> </script> <script defer="" src="https://c1.neweggimages.com/WebResource/Scripts/WWW/adapterScript~product~ProductDetail-da6079e7.js"> </script> <script async="" src="/KruBgE-3DRVqkyI6KQ/Szar2mwiL5/d04wdQ/LXxv/BQQHGDQ?v=035da9ea-c84c-3b78-e5b9-908bca3517dc" type="text/javascript"> </script> </body> </html>
问题原因
Newegg采用客户端JS动态渲染机制:服务器返回的初始HTML只是页面骨架,实际产品数据(包括价格)需要通过页面加载的JS脚本从后端接口获取,并动态渲染到页面中。用requests.get()只能拿到初始骨架,无法获取JS渲染后的内容。
解决方案
方法1:直接调用产品API(高效推荐)
Newegg提供产品详情API,可通过商品ID直接请求获取价格数据,无需解析HTML:
import requests # 替换为目标商品ID product_id = "N82E16814487547" api_url = f"https://www.newegg.com/api/product-detail?item={product_id}" # 添加请求头模拟浏览器访问,避免被拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(api_url, headers=headers) product_data = response.json() # 提取价格(API返回的结构可能会变动,需根据实际返回调整) current_price = product_data["ProductDetail"]["CurrentPrice"] print(f"GPU价格: ${current_price}")
方法2:用Selenium模拟浏览器渲染(适合需完整HTML场景)
通过Selenium启动真实浏览器,等待JS执行完成后再获取页面内容:
- 先安装依赖:
pip install selenium,并下载对应浏览器的驱动(如ChromeDriver,需与浏览器版本匹配) - 代码示例:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import time url = "https://www.newegg.com/evga-geforce-rtx-3080-ti-12g-p5-3967-kr/p/N82E16814487547" # 配置无头模式(不弹出浏览器窗口) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") # 启动浏览器并访问页面 driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 等待JS加载完成(时间可根据网络情况调整) time.sleep(3) # 获取渲染后的页面源码 page_source = driver.page_source driver.quit() # 解析HTML提取价格 doc = BeautifulSoup(page_source, "html.parser") price_element = doc.find(class_="price-current") if price_element: price_text = price_element.get_text(strip=True) print(f"GPU价格: {price_text}") else: print("未找到价格元素")
内容的提问来源于stack exchange,提问作者Zed
相关产品推荐
相关产品推荐

