使用Python3、BeautifulSoup4爬取supremenewyork.com UK遇文本混淆求助
Solutions for Scraping Supreme UK's Obfuscated Text
Hey there, I’ve tangled with Supreme’s anti-scraping measures a few times, so I can walk you through some practical fixes for this text obfuscation issue.
Option 1: Skip .text Entirely by Pulling Data Directly from Attributes or Raw Scripts
Supreme often hides key text in places that don’t rely on visible text nodes. Here are two reliable approaches:
- Extract from
data-*attributes: A lot of product details (names, prices, SKUs) are stored in customdata-attributes instead of rendered text. Use BeautifulSoup’sget()method to grab these directly:# Example for pulling a product name product_name = soup.find('div', class_='product-title').get('data-product-name') # Example for fetching price product_price = soup.find('span', class_='item-price').get('data-price-value') - Parse embedded JSON in script tags: Supreme frequently injects full product datasets into global JavaScript variables (like
window.__PRODUCTS__orwindow.preloadedData). Use regex to yank out this JSON string, then parse it with Python’sjsonmodule:import re import json # Get the raw HTML response (with your proxy setup) response = requests.get('https://www.supremenewyork.com/shop/all', proxies=your_proxy_config) # Match the JSON blob in the page's script tags json_pattern = re.compile(r'window\.preloadedData\s*=\s*(\{.*?\});') match = json_pattern.search(response.text) if match: product_data = json.loads(match.group(1)) # Access details directly from the parsed JSON for item in product_data['products']: print(f"Item: {item['name']} | Price: {item['price']}")
Option 2: Decode the Obfuscated Text
If the text is scrambled via encoding tricks or simple JS transformations, you can reverse the logic:
- Unescape HTML entities: Sometimes text is stored as HTML entities (like
Sfor "S"). Use Python’shtmlmodule to decode them:import html obfuscated_text = soup.find('div', class_='obfuscated-text').string decoded_text = html.unescape(obfuscated_text) - Reverse or rearrange character sequences: Some obfuscation reverses text or splits it into chunks. Check for patterns (e.g., reversed strings) and fix them directly:
obfuscated_text = 'emaN tcudorP' decoded_text = obfuscated_text[::-1] # Returns "Product Name" - Replicate JS decoding logic: If the site uses a custom JS function to unhide text, inspect the code via browser dev tools to figure out the steps, then mirror that logic in Python. For example, if JS uses Base64 decoding or a simple shift cipher, you can use Python’s
base64module or write a quick helper function to match it.
Quick Pro Tip
Always use your browser’s dev tools (Chrome/Firefox) to inspect the page’s HTML structure and network requests first. This will show you exactly where the real data lives—saving you from guessing at obfuscation methods unnecessarily.
内容的提问来源于stack exchange,提问作者user9510595
相关产品推荐
相关产品推荐

