You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Python+Selenium处理Grailed 500k URL的速度?

Optimizations for Grailed 0-Listings Scraper

1. Cut Per-Page Overhead (Fix Popup & Load Time Bottlenecks)

Block Popups at the Source

  • Pre-saved User Profiles: Instead of clicking popups every time, create a Chrome profile where you manually accept cookies and log in once. Load this profile in Selenium to retain session settings:
    from selenium.webdriver.chrome.options import Options
    
    options = Options()
    # Replace with your profile path (find via chrome://version/)
    options.add_argument(r"user-data-dir=C:\Users\YourUser\AppData\Local\Google\Chrome\User Data")
    options.add_argument("--profile-directory=Default")
    driver = webdriver.Chrome(options=options)
    
    Risk of ban is low here—this mimics a regular user's session. Test on a small batch first to confirm.
  • Network Interception: Block scripts that load cookie/login popups using Selenium's devtools API:
    driver.execute_cdp_cmd('Network.setBlockedURLs', {
        "urls": ["*cookie-banner.js", "*login-modal.js"]
    })
    driver.execute_cdp_cmd('Network.enable', {})
    
    Adjust the blocked URL patterns based on what you see in the browser's network tab.

Optimize Browser Load Settings

Disable non-essential resources to slash page load time:

options.add_argument("--headless=new")
options.add_argument("--blink-settings=imagesEnabled=false")
options.add_argument("--disable-css")
options.add_argument("--disable-extensions")
options.add_argument("--disable-plugins")

This can cut load time by 50% or more by skipping images, CSS, and extra extensions.

2. Concurrency & Scaling (Multi-Instance Strategy)

Parallel Processing Implementation

Use concurrent.futures to run multiple Selenium instances (each thread gets its own browser):

from concurrent.futures import ThreadPoolExecutor
from selenium import webdriver

def process_url(url):
    options = webdriver.ChromeOptions()
    # Add your optimized options here
    driver = webdriver.Chrome(options=options)
    try:
        driver.get(url)
        # Logic to check listing count
        listing_count = driver.find_element(By.CSS_SELECTOR, ".listings-count").text
        return (url, listing_count == "0")
    finally:
        driver.quit()

# Start with 5-10 instances, scale up gradually
with ThreadPoolExecutor(max_workers=8) as executor:
    results = executor.map(process_url, your_url_list)

Proxy Management

  • Why proxies are mandatory: Grailed will flag and ban repeated requests from a single IP. Use rotating residential proxies (less likely to be detected than datacenter proxies).
  • Integrate proxies into Selenium:
    proxy = "username:password@proxy-ip:port"
    options.add_argument(f"--proxy-server=http://{proxy}")
    
  • Cost breakdown: Residential proxies cost $50-$200/month for 100GB bandwidth. For 500k URLs, plan for ~100-200GB total.

Concurrency Limits

  • Start with 5-10 instances to test stability and avoid immediate bans.
  • Add small delays (1-2 seconds) between requests per instance to mimic human behavior:
    import time
    time.sleep(random.uniform(1, 2))
    

3. Alternative Tools & Faster Approaches

Switch to Playwright

Playwright is faster than Selenium, with built-in popup handling and better network control:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.context.clear_cookies()
    # Auto-accept cookies
    page.route("*/*", lambda route: route.continue_() if "cookie" not in route.request.url else route.abort())
    page.goto(url)
    listing_count = page.locator(".listings-count").text_content()
    browser.close()

Playwright can cut per-page time to 3-5 seconds, drastically reducing total runtime.

Use Undocumented API Endpoints

Inspect Grailed's network traffic (Chrome DevTools > Network tab) when loading a page. Look for API calls that return listing data (e.g., /api/shops/{shop_id}/listings). Use requests to call this endpoint directly:

import requests

headers = {
    "User-Agent": "Mozilla/5.0",
    "Cookie": "your-session-cookie"
}
response = requests.get("https://www.grailed.com/api/shops/123/listings", headers=headers)
data = response.json()
listing_count = data["total_count"]

This approach is 10-100x faster than browser rendering—total runtime could drop to hours instead of days.

4. Risk Mitigation & Anti-Ban Measures

  • Rotate User Agents: Use a list of real user agents and pick a random one for each instance.
  • Session Persistence: Keep browser sessions alive instead of restarting for every URL (reduces popup interactions and looks more human).
  • Error Handling: Detect 403 Forbidden or captcha pages, pause the instance, switch proxies, and retry after a delay.
  • Off-Peak Scraping: Scrape during low-traffic hours (e.g., 2 AM–8 AM UTC) to avoid triggering anti-scraping systems.

5. Clarifications on Your Original Ideas

  • Intercepting Popups: Saving a user profile is the safest way to bypass popups long-term. Network interception works but requires updating blocked URLs if Grailed changes their frontend.
  • Multi-Instance Cost:
    • Local: A machine with 16+ cores and 32+ GB RAM can run 10-20 instances for free.
    • Cloud: A t3.large AWS EC2 instance (2 cores, 8GB RAM) can run ~5 instances, costing ~$40/month. For 10 instances, budget $400/month plus proxy costs.

内容的提问来源于stack exchange,提问作者DudeNewToCoding

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 09:10:57