使用Scrapy-Splash抓取大型动态页面加载失败求助(Ubuntu环境)
Got it, let’s work through this problem together. You’re trying to scrape 24,453 dynamically loaded records from that Allgemeine Chirurgie page, but hitting roadblocks because loading all those entries at once is overwhelming your 4GB RAM setup—even with 50GB of swap. Here’s how to fix this without crashing your system:
1. Ditch Full Browser Loading—Use Headless Scraping or Direct API Calls
The biggest issue is that a regular browser will bloat its DOM with 24k+ elements, eating up all your RAM. Instead, go for a more efficient approach:
- Find the hidden API: Open your browser’s dev tools (F12), go to the Network tab, then click the "Mehr anzeigen" button. Look for XHR/fetch requests (usually to an endpoint like
/ajaxLoadMoreor similar). These requests are what the page uses to load new records—you can call this API directly, skipping the entire DOM. This is the fastest, lowest-memory method. - Headless browser with memory cleanup: If you can’t find the API, use tools like Puppeteer or Playwright in headless mode. After scraping each batch of 30 records, clear the page’s DOM or close/reopen the page to avoid memory buildup. For example, with Puppeteer:
const puppeteer = require('puppeteer'); (async () => { const browser = await puppeteer.launch(); let page = await browser.newPage(); await page.goto('https://www.sanego.de/Arzt/Allgemeine+Chirurgie/'); let totalScraped = 0; const totalRecords = 24453; while (totalScraped < totalRecords) { // Scrape current page's data here const records = await page.evaluate(() => { // Extract your data from the DOM elements return Array.from(document.querySelectorAll('.doctor-entry')).map(el => el.textContent.trim()); }); // Save records to file/database console.log(`Scraped ${records.length} records, total: ${totalScraped + records.length}`); totalScraped += records.length; // Click "Mehr anzeigen" await page.click('[title="Mehr anzeigen"]'); await page.waitForTimeout(1000); // Wait for load (adjust as needed) // Optional: Clear DOM to save memory await page.evaluate(() => document.body.innerHTML = ''); } await browser.close(); })();
2. Optimize Your Ubuntu System for Tight RAM
Your 4GB RAM is limited, so let’s squeeze every bit of performance out of it:
- Kill unnecessary background apps: Open
htopto see what’s eating memory—close unused browser tabs, IDEs, or media players. Freeing up 1-2GB will make a huge difference. - Limit browser memory (if you must use a GUI): If you need a visual browser, launch Chrome with a JS memory cap:
This limits the JS heap to 3GB, leaving 1GB for the system.google-chrome --js-flags="--max-old-space-size=3072" - Enable Zswap: Ubuntu 17.10 might not have this enabled by default. Zswap compresses swap data in RAM, reducing slow disk IO. Enable it temporarily:
To make it permanent, editecho 1 | sudo tee /sys/module/zswap/parameters/enabled/etc/default/grub, addzswap.enabled=1toGRUB_CMDLINE_LINUX, then runsudo update-gruband reboot.
3. Use a Lightweight Requests-Based Scraper
If you prefer Python, use requests to mimic the "Mehr anzeigen" requests directly—no browser needed. This uses almost no memory:
import requests import time BASE_URL = "https://www.sanego.de/Arzt/Allgemeine+Chirurgie/ajaxLoadMore" HEADERS = { "User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/117.0", "Referer": "https://www.sanego.de/Arzt/Allgemeine+Chirurgie/" } TOTAL_RECORDS = 24453 LIMIT = 30 offset = 0 with open('scraped_doctors.txt', 'a') as f: while offset < TOTAL_RECORDS: params = {"offset": offset, "limit": LIMIT} response = requests.post(BASE_URL, headers=HEADERS, params=params) # Parse the response (adjust based on actual response format) records = response.text # Replace with JSON parsing if response is JSON # Write to file f.write(records + "\n") offset += LIMIT print(f"Processed {offset}/{TOTAL_RECORDS} records") time.sleep(1) # Avoid hitting rate limits
4. Pro Tip: Avoid Rate Limits
Don’t forget to add delays between requests (like the time.sleep(1) above) to avoid getting your IP blocked. Some sites also check for consistent User-Agents, so make sure yours matches a real browser.
内容的提问来源于stack exchange,提问作者Ostap Didenko

