You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy-Splash抓取大型动态页面加载失败求助(Ubuntu环境)

Solution for Scraping 24k Dynamic Records on Ubuntu 17.10 (4GB RAM)

Got it, let’s work through this problem together. You’re trying to scrape 24,453 dynamically loaded records from that Allgemeine Chirurgie page, but hitting roadblocks because loading all those entries at once is overwhelming your 4GB RAM setup—even with 50GB of swap. Here’s how to fix this without crashing your system:

1. Ditch Full Browser Loading—Use Headless Scraping or Direct API Calls

The biggest issue is that a regular browser will bloat its DOM with 24k+ elements, eating up all your RAM. Instead, go for a more efficient approach:

  • Find the hidden API: Open your browser’s dev tools (F12), go to the Network tab, then click the "Mehr anzeigen" button. Look for XHR/fetch requests (usually to an endpoint like /ajaxLoadMore or similar). These requests are what the page uses to load new records—you can call this API directly, skipping the entire DOM. This is the fastest, lowest-memory method.
  • Headless browser with memory cleanup: If you can’t find the API, use tools like Puppeteer or Playwright in headless mode. After scraping each batch of 30 records, clear the page’s DOM or close/reopen the page to avoid memory buildup. For example, with Puppeteer:
    const puppeteer = require('puppeteer');
    
    (async () => {
      const browser = await puppeteer.launch();
      let page = await browser.newPage();
      await page.goto('https://www.sanego.de/Arzt/Allgemeine+Chirurgie/');
      
      let totalScraped = 0;
      const totalRecords = 24453;
    
      while (totalScraped < totalRecords) {
        // Scrape current page's data here
        const records = await page.evaluate(() => {
          // Extract your data from the DOM elements
          return Array.from(document.querySelectorAll('.doctor-entry')).map(el => el.textContent.trim());
        });
        
        // Save records to file/database
        console.log(`Scraped ${records.length} records, total: ${totalScraped + records.length}`);
        totalScraped += records.length;
    
        // Click "Mehr anzeigen"
        await page.click('[title="Mehr anzeigen"]');
        await page.waitForTimeout(1000); // Wait for load (adjust as needed)
    
        // Optional: Clear DOM to save memory
        await page.evaluate(() => document.body.innerHTML = '');
      }
    
      await browser.close();
    })();
    

2. Optimize Your Ubuntu System for Tight RAM

Your 4GB RAM is limited, so let’s squeeze every bit of performance out of it:

  • Kill unnecessary background apps: Open htop to see what’s eating memory—close unused browser tabs, IDEs, or media players. Freeing up 1-2GB will make a huge difference.
  • Limit browser memory (if you must use a GUI): If you need a visual browser, launch Chrome with a JS memory cap:
    google-chrome --js-flags="--max-old-space-size=3072"
    
    This limits the JS heap to 3GB, leaving 1GB for the system.
  • Enable Zswap: Ubuntu 17.10 might not have this enabled by default. Zswap compresses swap data in RAM, reducing slow disk IO. Enable it temporarily:
    echo 1 | sudo tee /sys/module/zswap/parameters/enabled
    
    To make it permanent, edit /etc/default/grub, add zswap.enabled=1 to GRUB_CMDLINE_LINUX, then run sudo update-grub and reboot.

3. Use a Lightweight Requests-Based Scraper

If you prefer Python, use requests to mimic the "Mehr anzeigen" requests directly—no browser needed. This uses almost no memory:

import requests
import time

BASE_URL = "https://www.sanego.de/Arzt/Allgemeine+Chirurgie/ajaxLoadMore"
HEADERS = {
    "User-Agent": "Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/117.0",
    "Referer": "https://www.sanego.de/Arzt/Allgemeine+Chirurgie/"
}
TOTAL_RECORDS = 24453
LIMIT = 30
offset = 0

with open('scraped_doctors.txt', 'a') as f:
    while offset < TOTAL_RECORDS:
        params = {"offset": offset, "limit": LIMIT}
        response = requests.post(BASE_URL, headers=HEADERS, params=params)
        
        # Parse the response (adjust based on actual response format)
        records = response.text  # Replace with JSON parsing if response is JSON
        
        # Write to file
        f.write(records + "\n")
        
        offset += LIMIT
        print(f"Processed {offset}/{TOTAL_RECORDS} records")
        time.sleep(1)  # Avoid hitting rate limits

4. Pro Tip: Avoid Rate Limits

Don’t forget to add delays between requests (like the time.sleep(1) above) to avoid getting your IP blocked. Some sites also check for consistent User-Agents, so make sure yours matches a real browser.

内容的提问来源于stack exchange,提问作者Ostap Didenko

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:17:09