You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python在不使用Selenium的情况下抓取任意JS渲染页面?

Got it, let's tackle this problem head-on. The key here is to bypass Selenium while still handling JavaScript-rendered content, and ideally have a solution that works across most such pages. Here are the most reliable approaches I've used in production:

1. Reverse-Engineer the Backend APIs (Most Efficient)

9 times out of 10, JavaScript-rendered pages are just pulling data from backend APIs via XHR or Fetch requests, then rendering it in the browser. Instead of scraping the final HTML, you can go straight to the source. Here's how:

  • Open your target page in Chrome/Firefox DevTools, switch to the Network tab, filter by XHR/Fetch.
  • Look for requests that return JSON data matching the content you see on the page (like product lists, article content).
  • Inspect the request headers, query parameters, and any authentication tokens (like Authorization headers) needed.
  • Use Python's requests library to mimic these requests directly.

Example code snippet:

import requests

# Mimic a real browser's headers to avoid being blocked
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'application/json, text/plain, */*'
}

# Replace with the actual API endpoint you found
response = requests.get(
    'https://api.example.com/api/v1/articles',
    headers=headers,
    params={'page': 1, 'limit': 10}
)

# Parse the JSON data directly
articles = response.json()
for article in articles:
    print(f"Title: {article['title']}, Content: {article['excerpt']}")

Pro tip: If the API uses dynamic parameters (like timestamps or signature tokens), you can often extract the JavaScript function that generates them from the page source, then run it in Python using execjs or js2py to replicate the parameters.

2. Use Lightweight Headless Browsers (Pyppeteer/Playwright)

If API reverse-engineering is too tricky (e.g., the site uses obfuscated APIs or heavy anti-scraping), lightweight headless browsers are a great alternative to Selenium. They launch a real browser in headless mode, let it render the page fully, then let you extract content.

I prefer Pyppeteer (Python wrapper for Google's Puppeteer) because it's lightweight and doesn't require the Selenium server setup. Here's how to use it:

import asyncio
from pyppeteer import launch

async def scrape_rendered_page(url):
    # Launch headless Chrome
    browser = await launch(headless=True, args=['--no-sandbox'])
    page = await browser.newPage()
    
    # Set a realistic user agent
    await page.setUserAgent('Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')
    
    # Wait for the page to fully load (network idle means no new requests for 500ms)
    await page.goto(url, waitUntil='networkidle2')
    
    # Extract content: either get full HTML or query specific elements
    full_html = await page.content()
    page_title = await page.title()
    
    # Extract text from all product cards using CSS selectors
    product_names = await page.querySelectorAllEval(
        '.product-card__title',
        '(nodes) => nodes.map(node => node.textContent.trim())'
    )
    
    await browser.close()
    return full_html, page_title, product_names

# Run the async function
result = asyncio.get_event_loop().run_until_complete(
    scrape_rendered_page('https://example.com/js-heavy-store')
)
print(f"Page Title: {result[1]}")
print(f"Products: {result[2]}")

Note: You'll need Node.js installed for Pyppeteer to work, but it's a one-time setup. For even more control, you can use Playwright (Microsoft's alternative to Puppeteer) which has official Python bindings too.

3. Execute Inline JavaScript with js2py

For pages where content is generated by simple inline JavaScript (no browser API dependencies like document or window), you can extract the relevant JS code and run it directly in Python using js2py. This avoids spinning up a browser entirely.

Example: Suppose a page uses a JS function to calculate dynamic prices:

import js2py

# Extract the JavaScript function from the page source
js_function = """
function calculateFinalPrice(basePrice, taxRate, discount) {
    let taxed = basePrice * (1 + taxRate/100);
    return taxed * (1 - discount/100);
}
"""

# Execute the JS code in a Python context
js_context = js2py.EvalJs()
js_context.execute(js_function)

# Call the JS function from Python
final_price = js_context.calculateFinalPrice(100, 8, 15)
print(f"Final Price: ${final_price:.2f}")  # Output: $91.80

This works best for isolated JS logic—if the code relies on browser-specific objects, you'll need a headless browser instead.

4. Build a Fallback Pipeline for Universal Compatibility

There's no one-size-fits-all solution, but you can build a pipeline that tries approaches in order of efficiency:

  1. First, attempt to reverse-engineer the API (fastest, lowest resource usage).
  2. If that fails, use a headless browser to render the page.
  3. For edge cases with simple inline JS, use js2py to supplement.

Also, don't forget anti-scraping best practices:

  • Add random delays between requests (time.sleep(random.uniform(1, 3))).
  • Rotate user agents and proxies if the site blocks repeated requests.
  • Handle cookies and sessions properly to mimic a real user.

内容的提问来源于stack exchange,提问作者jwoojin9

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:44:37