如何用Python在不使用Selenium的情况下抓取任意JS渲染页面?
Got it, let's tackle this problem head-on. The key here is to bypass Selenium while still handling JavaScript-rendered content, and ideally have a solution that works across most such pages. Here are the most reliable approaches I've used in production:
9 times out of 10, JavaScript-rendered pages are just pulling data from backend APIs via XHR or Fetch requests, then rendering it in the browser. Instead of scraping the final HTML, you can go straight to the source. Here's how:
- Open your target page in Chrome/Firefox DevTools, switch to the Network tab, filter by XHR/Fetch.
- Look for requests that return JSON data matching the content you see on the page (like product lists, article content).
- Inspect the request headers, query parameters, and any authentication tokens (like
Authorizationheaders) needed. - Use Python's
requestslibrary to mimic these requests directly.
Example code snippet:
import requests # Mimic a real browser's headers to avoid being blocked headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'application/json, text/plain, */*' } # Replace with the actual API endpoint you found response = requests.get( 'https://api.example.com/api/v1/articles', headers=headers, params={'page': 1, 'limit': 10} ) # Parse the JSON data directly articles = response.json() for article in articles: print(f"Title: {article['title']}, Content: {article['excerpt']}")
Pro tip: If the API uses dynamic parameters (like timestamps or signature tokens), you can often extract the JavaScript function that generates them from the page source, then run it in Python using execjs or js2py to replicate the parameters.
If API reverse-engineering is too tricky (e.g., the site uses obfuscated APIs or heavy anti-scraping), lightweight headless browsers are a great alternative to Selenium. They launch a real browser in headless mode, let it render the page fully, then let you extract content.
I prefer Pyppeteer (Python wrapper for Google's Puppeteer) because it's lightweight and doesn't require the Selenium server setup. Here's how to use it:
import asyncio from pyppeteer import launch async def scrape_rendered_page(url): # Launch headless Chrome browser = await launch(headless=True, args=['--no-sandbox']) page = await browser.newPage() # Set a realistic user agent await page.setUserAgent('Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') # Wait for the page to fully load (network idle means no new requests for 500ms) await page.goto(url, waitUntil='networkidle2') # Extract content: either get full HTML or query specific elements full_html = await page.content() page_title = await page.title() # Extract text from all product cards using CSS selectors product_names = await page.querySelectorAllEval( '.product-card__title', '(nodes) => nodes.map(node => node.textContent.trim())' ) await browser.close() return full_html, page_title, product_names # Run the async function result = asyncio.get_event_loop().run_until_complete( scrape_rendered_page('https://example.com/js-heavy-store') ) print(f"Page Title: {result[1]}") print(f"Products: {result[2]}")
Note: You'll need Node.js installed for Pyppeteer to work, but it's a one-time setup. For even more control, you can use Playwright (Microsoft's alternative to Puppeteer) which has official Python bindings too.
For pages where content is generated by simple inline JavaScript (no browser API dependencies like document or window), you can extract the relevant JS code and run it directly in Python using js2py. This avoids spinning up a browser entirely.
Example: Suppose a page uses a JS function to calculate dynamic prices:
import js2py # Extract the JavaScript function from the page source js_function = """ function calculateFinalPrice(basePrice, taxRate, discount) { let taxed = basePrice * (1 + taxRate/100); return taxed * (1 - discount/100); } """ # Execute the JS code in a Python context js_context = js2py.EvalJs() js_context.execute(js_function) # Call the JS function from Python final_price = js_context.calculateFinalPrice(100, 8, 15) print(f"Final Price: ${final_price:.2f}") # Output: $91.80
This works best for isolated JS logic—if the code relies on browser-specific objects, you'll need a headless browser instead.
There's no one-size-fits-all solution, but you can build a pipeline that tries approaches in order of efficiency:
- First, attempt to reverse-engineer the API (fastest, lowest resource usage).
- If that fails, use a headless browser to render the page.
- For edge cases with simple inline JS, use js2py to supplement.
Also, don't forget anti-scraping best practices:
- Add random delays between requests (
time.sleep(random.uniform(1, 3))). - Rotate user agents and proxies if the site blocks repeated requests.
- Handle cookies and sessions properly to mimic a real user.
内容的提问来源于stack exchange,提问作者jwoojin9

