使用BS4抓取网站隐藏变量:无法获取HM商品尺码及库存状态
Hey there! Let’s work through this issue you’re having with scraping size availability from that HM product page. I’ve run into this exact problem with fashion retail sites before—they rarely serve up size/stock data directly in visible HTML tags, which is why you can’t find it even after checking the full source.
Why You Can’t See the Data in Static HTML
HM loads product details like sizes and stock status dynamically using JavaScript. Instead of embedding this info in <select> or <option> tags upfront, they store it in a hidden JSON object within a <script> tag on the page. When the page loads, the frontend JS pulls this data to build the size dropdown.
Step-by-Step Solution
Here’s how to extract that hidden data using Python, Requests, and BeautifulSoup:
Fetch the page with proper headers
Make sure you send a validUser-Agentto avoid being blocked by HM’s anti-scraping measures.Locate the script tag with product data
Search through all<script>tags on the page for a global variable (usually something likewindow.product) that contains the full product dataset.Extract and parse the JSON
Use regex to pull the JSON string from the script content, then convert it to a Python dictionary to access sizes and stock status.
Example Code
import requests from bs4 import BeautifulSoup import json import re # Target product URL url = "http://www2.hm.com/en_in/productpage.0648256001.html" # Mimic a real browser request headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # Get the page content response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # Hunt for the script with product data for script in soup.find_all("script"): script_text = script.string # Check if the script contains product and size data if script_text and "window.product" in script_text and "sizes" in script_text: # Extract the JSON object using regex json_match = re.search(r'window\.product\s*=\s*({.*?});', script_text, re.DOTALL) if json_match: product_json = json_match.group(1) product_data = json.loads(product_json) # Extract size and availability info size_list = product_data.get("sizes", []) print("Size Availability:") for size in size_list: size_label = size.get("name") in_stock = size.get("available", False) print(f"- {size_label}: {'In Stock' if in_stock else 'Sold Out'}") break
Key Notes
- Anti-scraping checks: If you get blocked, try rotating your
User-Agentor adding a small delay between requests. - Variable names: Sometimes HM uses different variable names (like
initialProductData), so adjust the regex if needed. Just search the script content for keywords like "sizes" or "availability" to find the right variable.
内容的提问来源于stack exchange,提问作者mehul mittal

