Beautiful Soup返回不完整HTML脚本:无法定位Blue Nile钻石列表元素
Hey there! Let's break down why your current code isn't picking up the diamond elements and how to fix it.
问题根源
The issue here is that Blue Nile loads its diamond list dynamically using JavaScript. When you use urllib.request to fetch the page, you're only getting the raw static HTML sent by the server—this doesn't include the diamond elements, because those are rendered after the page loads by pulling data from an API and injecting it into the DOM. That's why your diamonds array comes back empty, even though you can see the catalog-view-offer-wrapper classes in your browser's DevTools (which shows the fully rendered DOM).
解决方案1:使用Selenium模拟浏览器渲染
Selenium launches a real browser (like Chrome) that loads and executes JavaScript, so it can capture the fully rendered page. Here's how to adjust your code:
步骤1:安装依赖
First, install Selenium and download the appropriate browser driver (e.g., ChromeDriver for Chrome):
pip install selenium
Make sure the driver is in your system PATH or specify its path in the code.
步骤2:修改后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup as soup my_url = 'https://www.bluenile.com/uk/diamond-search?tag=none&track=NavDiaVAll' # Initialize Chrome browser (adjust driver path if needed) driver = webdriver.Chrome() driver.get(my_url) # Wait up to 10 seconds for the diamond elements to load try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "catalog-view-offer-wrapper")) ) finally: # Get the fully rendered page HTML page_html = driver.page_source driver.quit() # Parse the rendered HTML page_soup = soup(page_html, "html.parser") diamonds = page_soup.findAll("div", {"class": "catalog-view-offer-wrapper"}) print(len(diamonds)) # This should now return the correct count!
解决方案2:直接调用API(更高效)
Instead of simulating a browser, you can find the API that Blue Nile uses to fetch diamond data and call it directly. This is faster and more reliable than scraping rendered HTML.
步骤1:找到API接口
- Open your browser's DevTools (F12)
- Go to the Network tab
- Refresh the Blue Nile diamond page
- Look for XHR/fetch requests (filter by "XHR")—you'll see a request to an endpoint like
https://www.bluenile.com/api/public/diamond-search/v2/searchwith query parameters for country, currency, page size, etc.
步骤2:API请求代码示例
import requests import csv # Example API endpoint (adjust parameters like page/pageSize as needed) api_url = "https://www.bluenile.com/api/public/diamond-search/v2/search?country=GB&language=en¤cy=GBP&productSet=blue_nile&shape=RD&page=1&pageSize=100" # Add headers to mimic a real browser request headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # Fetch the data response = requests.get(api_url, headers=headers) data = response.json() # Extract diamond details diamonds = data.get('diamonds', []) print(f"Found {len(diamonds)} diamonds") # Save to CSV with open('bluenile_diamonds.csv', 'w', newline='', encoding='utf-8') as csv_file: # Define CSV columns (adjust based on the data you want) fieldnames = ['id', 'carat', 'cut', 'color', 'clarity', 'price_gbp'] writer = csv.DictWriter(csv_file, fieldnames=fieldnames) writer.writeheader() for diamond in diamonds: writer.writerow({ 'id': diamond['id'], 'carat': diamond['carat'], 'cut': diamond['cut'], 'color': diamond['color'], 'clarity': diamond['clarity'], 'price_gbp': diamond['price']['amount'] })
注意事项
- Always check a website's
robots.txtand Terms of Service before scraping—avoid making too many requests in a short time to prevent getting blocked. - API endpoints might change over time, so you may need to recheck the Network tab if the code stops working.
内容的提问来源于stack exchange,提问作者Sehgal123

