You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests编写的Python web crawler无法获取完整页面,如何解决?

Fixing Incomplete Page Issues with Async Content in Python Web Crawlers

Hey Kevin, totally get where you're coming from—dealing with async-loaded content can be super frustrating when you're just starting out with web crawlers! Let's break down why your requests-based spider isn't grabbing the full page, and walk through actionable fixes.

Why requests Isn't Working

The requests library only fetches the initial static HTML sent by the server. If the site uses asynchronous loading (like AJAX, Fetch API, or client-side JavaScript rendering), the content you want is loaded after the initial page loads—when your browser runs the site's JavaScript. requests doesn't execute JS, so it can't see that dynamic content.

Solution 1: Target the Direct API Endpoints (Most Efficient!)

Many sites load async content by making background API calls. You can skip rendering the page entirely by calling these APIs directly:

  • Open your browser's DevTools (press F12) and go to the Network tab.
  • Refresh the page, then filter requests by XHR or Fetch (these are the calls that load dynamic content).
  • Look for requests that return the data you need (usually in JSON format). Check the request URL, headers, and any parameters.
  • Replicate that request with requests in your code. Make sure to include necessary headers (like User-Agent, Referer, or even cookies if the site requires authentication).

Example code snippet:

import requests

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Referer": "https://example.com/target-page"
}

api_url = "https://example.com/api/get-dynamic-content"
response = requests.get(api_url, headers=headers)
data = response.json()

# Now you can parse the JSON data directly!
print(data)

Solution 2: Use a Headless Browser to Render JavaScript

If the site's dynamic content is tied to complex JS rendering (or you can't find the API endpoints), use a headless browser tool that executes JS just like a real browser. Two popular options are:

Playwright is lightweight, fast, and supports all major browsers.

  1. Install it first:
    pip install playwright
    playwright install  # This installs the browser binaries
    
  2. Example code to load and get full page content:
    from playwright.sync_api import sync_playwright
    
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)  # Headless means no visible window
        page = browser.new_page()
        page.goto("https://example.com/target-page")
        
        # Wait for a specific element that loads dynamically (adjust selector as needed)
        page.wait_for_selector(".dynamic-content-class")
        
        # Get the full rendered HTML
        full_html = page.content()
        print(full_html)
        
        browser.close()
    

Selenium (Older but Widely Used)

Selenium is another option, though it requires downloading a browser driver (e.g., ChromeDriver for Chrome).

  1. Install:
    pip install selenium
    
  2. Example code:
    from selenium import webdriver
    from selenium.webdriver.common.by import By
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    
    # Initialize Chrome driver (make sure ChromeDriver is in your PATH)
    driver = webdriver.Chrome()
    driver.get("https://example.com/target-page")
    
    # Wait for dynamic content to load
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "dynamic-content-class"))
    )
    
    # Get full page source
    full_html = driver.page_source
    print(full_html)
    
    driver.quit()
    

Quick Tips to Avoid Headaches

  • Mimic a real user: Always include a valid User-Agent header to avoid being blocked by anti-scraping tools.
  • Handle scroll loading: If content loads when scrolling, use page.mouse.wheel() (Playwright) or driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") (Selenium) to trigger loading.
  • Check robots.txt: Make sure the site allows crawling the content you're targeting (look at https://example.com/robots.txt).

内容的提问来源于stack exchange,提问作者Kevin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:34:29