You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3网页爬取:如何延长请求时长及Flipkart爬取报错求助

Hey Prateek, let’s break down your issues step by step—first the Flipkart blocking problem, then adjusting request timing in your Python crawler.

Fixing Flipkart’s "Something is Wrong!!!" Block

That error after 14 pages is almost certainly Flipkart’s anti-bot system flagging your crawler. Here’s how to get around it:

  • Slow down with random delays: Rapid, uniform requests are a dead giveaway for bots. Replace fixed pauses with random wait times to mimic human browsing:

    import time
    import random
    
    # Add this after navigating to each new page
    time.sleep(random.uniform(2, 6))  # Wait between 2-6 seconds randomly
    
  • Rotate User-Agents: Flipkart tracks consistent User-Agent strings. Use a list of real browser identifiers and pick one randomly for each session:

    from selenium.webdriver.chrome.options import Options
    
    user_agents = [
        "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
        "Mozilla/5.0 (Macintosh; Intel Mac OS X 14_0) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.0 Safari/605.1.15",
        # Add more real User-Agents from browser dev tools here
    ]
    
    chrome_options = Options()
    chrome_options.add_argument(f"user-agent={random.choice(user_agents)}")
    driver = webdriver.Chrome(options=chrome_options)
    
  • Use proxy IP rotation: After 14 pages, your IP is likely flagged. Rotate proxies to avoid bans (use reputable proxy services for reliability):

    proxies = ["192.168.1.1:8080", "10.0.0.1:3128"]  # Replace with real proxies
    proxy = random.choice(proxies)
    chrome_options.add_argument(f"--proxy-server={proxy}")
    
  • Simulate human interactions: Don’t just load pages—scroll, hover, or click minor elements to look less robotic:

    # Scroll to bottom of the page to trigger lazy-loaded content
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(random.uniform(1, 3))  # Wait after scrolling
    
  • Maintain session consistency: Avoid reinitializing the Selenium driver every page. Keep the same session alive to mimic a real user’s browsing session.

Increasing Request/Page Load Wait Times in Python 3

The approach depends on whether you’re using Selenium (for dynamic pages) or requests/bs4 (for static pages):

For Selenium (Dynamic Pages)

  • Explicit Waits (Recommended): Wait for specific elements to load before proceeding—this is more efficient than fixed sleeps because it stops waiting as soon as the element is ready:

    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    from selenium.webdriver.common.by import By
    
    # Wait up to 15 seconds for Flipkart's comment section to load
    wait = WebDriverWait(driver, 15)
    comment_section = wait.until(EC.presence_of_element_located((By.CLASS_NAME, "_16Bl6k")))
    
  • Implicit Waits: Set a global wait time for all element searches—if an element isn’t found immediately, the driver waits up to the specified time:

    driver.implicitly_wait(10)  # Wait up to 10 seconds for elements to appear
    
  • Fixed Sleeps: Use time.sleep() for simple pauses, but always pair with randomness to avoid detection.

For Requests/BeautifulSoup (Static Pages)

  • Adjust request timeout: When making a requests.get() call, set the timeout parameter to extend how long the request waits for a server response:

    import requests
    
    response = requests.get(url, timeout=30)  # Wait up to 30 seconds for the server to respond
    
  • Add delays between requests: Again, use time.sleep(random.uniform(2,5)) between requests to slow down your crawl rate.

内容的提问来源于stack exchange,提问作者Prateek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:09:47