如何用Scrapy+Selenium抓取动态加载内容并提取文本入库?
Hey there! Let's work through your dynamic content scraping problem together—you're already halfway there with your existing setup, so let's fix the text extraction gap and explore other reliable options.
1. Selenium Text Extraction Example
Your current code is capturing screenshots, but we can tweak it to pull actual page text by leveraging Selenium's access to the fully rendered DOM. Here's how to adjust your spider to extract clean, usable text:
import scrapy from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException import w3lib.html class QuotesSpider(scrapy.Spider): name = "quotes" def __init__(self): # Initialize Chrome driver (ensure chromedriver is in your system PATH) self.driver = webdriver.Chrome() self.wait = WebDriverWait(self.driver, 15) # Longer timeout for slow-loading dynamic content def start_requests(self): # Replace this with URLs loaded from your Excel file urls = [ 'https://www.analog.com/en/products/landing-pages/new-products-listing.html', ] for url in urls: yield scrapy.Request(url=url, callback=self.parse_with_selenium) def parse_with_selenium(self, response): self.driver.get(response.url) try: # Wait for a critical dynamic element to load (update the selector to match your target page) self.wait.until( EC.presence_of_element_located((By.CSS_SELECTOR, ".product-listing-container")) ) # Optional: Scroll to bottom to trigger lazy-loaded content self.driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait extra for lazy content to finish loading self.wait.until( EC.presence_of_element_located((By.CSS_SELECTOR, ".lazy-loaded-item")) ) # Extract and clean the page text page_source = self.driver.page_source # Strip HTML tags and normalize whitespace cleaned_text = w3lib.html.remove_tags(page_source) cleaned_text = " ".join(cleaned_text.split()) # Pass the cleaned text to your existing database save logic self.save_to_db(response.url, cleaned_text) except TimeoutException: self.logger.error(f"Timeout waiting for dynamic content on {response.url}") # Note: Reuse the driver for multiple URLs instead of quitting immediately for efficiency def save_to_db(self, url, text): # Insert your existing database insertion code here self.logger.info(f"Successfully saved content for {url}")
Key tweaks here:
- Directly use Selenium's driver to load pages and wait for critical dynamic elements
- Extract the full rendered page source, then clean it to plain text
- Handle lazy-loaded content with scrolling and targeted waits
- Integrates seamlessly with your existing database workflow
2. Alternative Solutions for Dynamic Content Scraping
a. Playwright (Modern, Low-Boilerplate Alternative to Selenium)
Playwright is more reliable for dynamic pages and requires less setup than Selenium. Here's a quick snippet to integrate it with Scrapy:
from playwright.sync_api import sync_playwright import scrapy import w3lib.html class PlaywrightSpider(scrapy.Spider): name = "playwright_quotes" def start_requests(self): urls = ['https://www.analog.com/en/products/landing-pages/new-products-listing.html'] with sync_playwright() as p: browser = p.chromium.launch(headless=True) # Headless mode saves resources page = browser.new_page() for url in urls: # Wait until network is idle to ensure all content loads page.goto(url, wait_until="networkidle") # Trigger lazy loading by scrolling page.evaluate("window.scrollTo(0, document.body.scrollHeight);") page.wait_for_load_state("networkidle") # Extract and clean text page_source = page.content() cleaned_text = w3lib.html.remove_tags(page_source) cleaned_text = " ".join(cleaned_text.split()) # Save to your database here self.logger.info(f"Extracted text for {url}") browser.close()
b. Scrape Underlying API Directly
Most dynamic pages load content via AJAX/API calls. Use your browser's DevTools (Network tab, filter for XHR/Fetch) to find these endpoints. Calling the API directly is faster and more reliable than scraping the rendered page:
- Open the target page in Chrome/Firefox DevTools
- Go to the Network tab, refresh the page, and look for requests returning JSON data
- Copy the API URL and use Scrapy's
Requestto fetch the JSON, then extract text directly from the response
c. Improve Splash Usage
If you want to stick with Splash, adjust your Lua script to wait longer and ensure all content loads:
from scrapy_splash import SplashRequest # Inside your spider's start_requests method yield SplashRequest( url=url, callback=self.parse, endpoint='execute', args={ 'lua_source': ''' function main(splash, args) splash:go(args.url) splash:wait(5) # Wait for initial page load splash:runjs("window.scrollTo(0, document.body.scrollHeight);") splash:wait(3) # Wait for lazy content to load return splash:html() # Return fully rendered HTML end ''' } )
Final Tips
- Always respect the website's
robots.txtand terms of service - For large batches, use headless browsers to reduce resource usage
- Add retry logic for failed requests to handle intermittent loading issues
内容的提问来源于stack exchange,提问作者Amarjot Singh

