使用BeautifulSoup爬取阿迪达斯男鞋页面无法获取价格的优化方案
Hey Greg, I’ve run into this exact issue with dynamic e-commerce sites before—here’s why your current code is missing prices and how to fix it properly for scaling to hundreds of shoes:
Why Your Current Code Isn’t Getting Prices
Adidas loads product data (including prices) dynamically with JavaScript after the initial static HTML loads. When you use requests.get() to fetch the page, you’re only getting the empty shell of the page before the JS fills in the prices. That’s why your results have empty price fields!
Best Solution: Use Adidas’ Official API (Fast & Scalable)
E-commerce sites like Adidas almost always use backend APIs to serve product data to the frontend. By calling this API directly, you’ll get full, structured data (including real-time prices) without needing to parse messy HTML. This is way faster and more reliable for scraping hundreds of products.
Step-by-Step Optimized Code
import requests def fetch_adidas_men_shoes(max_pages=3): # Adidas product search API endpoint (found via browser dev tools) api_url = "https://www.adidas.com/api/search/product" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/84.0.4147.105 Safari/537.36", "Referer": "https://www.adidas.com/us/men-shoes" } all_products = [] offset = 0 items_per_page = 48 # Default Adidas page size for page in range(max_pages): params = { "sitePath": "us", "query": "men-shoes", "offset": offset, "limit": items_per_page } # Send API request response = requests.get(api_url, headers=headers, params=params, timeout=15) if response.status_code != 200: print(f"Failed to load page {page+1}: Status code {response.status_code}") break # Parse JSON response product_data = response.json() raw_products = product_data.get("rawProducts", []) if not raw_products: print("No more products to fetch") break # Extract key details for product in raw_products: product_info = { "name": product.get("name", "Unnamed Product"), "current_price": product.get("price", {}).get("current", "N/A"), "original_price": product.get("price", {}).get("original", "N/A"), "product_url": f"https://www.adidas.com/us/{product.get('slug')}.html", "image_url": product.get("image", {}).get("url") } all_products.append(product_info) offset += items_per_page return all_products # Example usage: Fetch 3 pages (~144 products) shoes = fetch_adidas_men_shoes(max_pages=3) for shoe in shoes[:5]: print(f"Name: {shoe['name']}") print(f"Current Price: ${shoe['current_price']}") print(f"Original Price: ${shoe['original_price'] if shoe['original_price'] != 'N/A' else 'No original price'}") print(f"Link: {shoe['product_url']}\n")
Why This Works Better
- Full, Real-Time Data: The API returns complete product details, including prices that were missing from static HTML.
- Scalable: Adjust
max_pagesto fetch as many products as you need (hundreds or more). - Faster: JSON parsing is way quicker than HTML parsing with BeautifulSoup, so you’ll get results in a fraction of the time.
- Lower Risk of Blocking: API requests are more lightweight than browser emulation, reducing the chance of being flagged as a bot (still, add delays between requests if needed!).
Alternative: Use Selenium for Browser Emulation
If you absolutely need to use BeautifulSoup (e.g., for other page elements), you can use Selenium to wait for the JS to load the prices. Note this is slower and less scalable for hundreds of products:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from bs4 import BeautifulSoup def scrape_with_selenium(): url = "https://www.adidas.com/us/men-shoes" # Configure Chrome in headless mode (no visible window) chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/84.0.4147.105 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) driver.get(url) # Wait up to 10 seconds for prices to load try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "gl-price")) ) except Exception as e: print(f"Error waiting for prices: {e}") driver.quit() return [] # Parse the fully loaded page with BeautifulSoup soup = BeautifulSoup(driver.page_source, "html.parser") driver.quit() products = soup.find_all("div", class_="gl-product-card color-variations__fixed-size glass-product-card___17N3p") shoe_list = [] for product in products: name = product.find("div", class_="gl-product-card__name").text.strip() if product.find("div", class_="gl-product-card__name") else "N/A" price = product.find("div", class_="gl-price").text.strip() if product.find("div", class_="gl-price") else "N/A" shoe_list.append({"name": name, "price": price}) return shoe_list # Example usage selenium_shoes = scrape_with_selenium() for shoe in selenium_shoes[:5]: print(f"{shoe['name']} - {shoe['price']}")
Final Notes
- Always check Adidas’
robots.txtand terms of service before scraping. - For large-scale scraping, add random delays between requests (e.g.,
time.sleep(random.uniform(1,3))) to avoid being blocked. - The API endpoint might change over time—if it stops working, use your browser’s dev tools (Network tab > XHR) to find the latest product search endpoint.
内容的提问来源于stack exchange,提问作者Greg

