You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python3、BeautifulSoup4爬取supremenewyork.com UK遇文本混淆求助

Solutions for Scraping Supreme UK's Obfuscated Text

Hey there, I’ve tangled with Supreme’s anti-scraping measures a few times, so I can walk you through some practical fixes for this text obfuscation issue.

Option 1: Skip .text Entirely by Pulling Data Directly from Attributes or Raw Scripts

Supreme often hides key text in places that don’t rely on visible text nodes. Here are two reliable approaches:

  • Extract from data-* attributes: A lot of product details (names, prices, SKUs) are stored in custom data- attributes instead of rendered text. Use BeautifulSoup’s get() method to grab these directly:
    # Example for pulling a product name
    product_name = soup.find('div', class_='product-title').get('data-product-name')
    # Example for fetching price
    product_price = soup.find('span', class_='item-price').get('data-price-value')
    
  • Parse embedded JSON in script tags: Supreme frequently injects full product datasets into global JavaScript variables (like window.__PRODUCTS__ or window.preloadedData). Use regex to yank out this JSON string, then parse it with Python’s json module:
    import re
    import json
    
    # Get the raw HTML response (with your proxy setup)
    response = requests.get('https://www.supremenewyork.com/shop/all', proxies=your_proxy_config)
    # Match the JSON blob in the page's script tags
    json_pattern = re.compile(r'window\.preloadedData\s*=\s*(\{.*?\});')
    match = json_pattern.search(response.text)
    if match:
        product_data = json.loads(match.group(1))
        # Access details directly from the parsed JSON
        for item in product_data['products']:
            print(f"Item: {item['name']} | Price: {item['price']}")
    

Option 2: Decode the Obfuscated Text

If the text is scrambled via encoding tricks or simple JS transformations, you can reverse the logic:

  • Unescape HTML entities: Sometimes text is stored as HTML entities (like S for "S"). Use Python’s html module to decode them:
    import html
    
    obfuscated_text = soup.find('div', class_='obfuscated-text').string
    decoded_text = html.unescape(obfuscated_text)
    
  • Reverse or rearrange character sequences: Some obfuscation reverses text or splits it into chunks. Check for patterns (e.g., reversed strings) and fix them directly:
    obfuscated_text = 'emaN tcudorP'
    decoded_text = obfuscated_text[::-1]  # Returns "Product Name"
    
  • Replicate JS decoding logic: If the site uses a custom JS function to unhide text, inspect the code via browser dev tools to figure out the steps, then mirror that logic in Python. For example, if JS uses Base64 decoding or a simple shift cipher, you can use Python’s base64 module or write a quick helper function to match it.

Quick Pro Tip

Always use your browser’s dev tools (Chrome/Firefox) to inspect the page’s HTML structure and network requests first. This will show you exactly where the real data lives—saving you from guessing at obfuscation methods unnecessarily.

内容的提问来源于stack exchange,提问作者user9510595

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:08:19