求助:如何在Scrapy爬虫中结合Selenium点击Kickstarter的see more!按钮
Fixing Your Kickstarter Spider: Clicking "See More!" & Scraping All Backers
Hey there! Let's work through the issues in your code so you can get that "see more!" button working and scrape all the backer data. Here's what's going wrong and how to fix it:
Key Issues in Your Current Code
- Incorrect constructor: You wrote
_init_instead of__init__(double underscores on both sides), which means your driver isn't being initialized properly. - Typo in driver reference:
selfdrivershould beself.driver(missing the dot), which would throw an AttributeError. - No wait for elements: The "see more!" button might not be loaded yet when you try to click it—you need to wait for it to be clickable.
- Misaligned parsing logic: You're trying to use the original Scrapy
responseinstead of the page source from Selenium, and yourparse2method doesn't have access to thebackersvariable. - Loop logic flaw: You're clicking the button but not parsing the data between clicks, and you don't check if new data is actually loaded before stopping.
Corrected Code
Here's a revised version of your spider that fixes these issues:
from scrapy import Spider from scrapy.http import HtmlResponse from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException class KickstarterSpider(Spider): name = 'kickstarter' allowed_domains = ['kickstarter.com'] start_urls = [ "https://www.kickstarter.com/projects/zwim/zwim-smart-swimming-goggles/community", "https://www.kickstarter.com/projects/zunik/oriboard-the-amazing-origami-multifunctional-cutti/community" ] def __init__(self): # Initialize Chrome driver (ensure chromedriver is in your PATH) self.driver = webdriver.Chrome() self.wait = WebDriverWait(self.driver, 10) # 10-second wait for dynamic elements def parse(self, response): # Load target page with Selenium self.driver.get(response.url) seen_backers = set() # Track collected backers to avoid duplicates while True: # Wait for backer elements to load, then parse with Scrapy self.wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '.founding-backer.community-block-content'))) page_source = self.driver.page_source selenium_response = HtmlResponse(url=response.url, body=page_source, encoding='utf-8') # Extract and process current page's backers backers = selenium_response.css('.founding-backer.community-block-content') for b in backers: name = b.css('.name.js-founding-backer-name::text').get().strip() backed = b.css('.backing-count.js-founding-backer-backings::text').get().strip() # Check for duplicates to stop scraping if (name, backed) not in seen_backers: seen_backers.add((name, backed)) yield { 'backer_name': name, 'projects_backed': backed } else: # No new data left, clean up and exit self.driver.quit() return # Try to click the "see more!" button try: see_more_btn = self.wait.until(EC.element_to_be_clickable((By.XPATH, '//button[contains(text(), "see more!")]'))) see_more_btn.click() # Wait for new content to load (wait until last old backer element is gone) self.wait.until(EC.staleness_of(backers[-1])) except TimeoutException: # No more button or new data, exit loop self.driver.quit() return
What Changed & Why
- Proper Driver Setup: The
__init__method uses double underscores and initializes aWebDriverWaitto handle dynamic content loading gracefully. - Wait for Elements: We wait until backers are visible and the "see more!" button is clickable—this avoids errors from interacting with unloaded elements.
- Scrape Between Clicks: After each page load (and after clicking the button), we parse the current backer data and track duplicates. When we hit a duplicate, we stop (since we've collected all unique data).
- Selenium + Scrapy Integration: We convert Selenium's page source into a Scrapy
HtmlResponse, so you can use Scrapy's familiar CSS selectors for parsing. - Clean Exit: The driver quits automatically when scraping is done, preventing leftover browser windows.
Quick Tips
- If
chromedriverisn't in your system PATH, specify its location withwebdriver.Chrome(executable_path='/your/path/to/chromedriver'). - If the button's selector changes, update the XPath to match the actual button text or attributes (using
contains(text(), "see more!")is more robust than a fixed ID).
内容的提问来源于stack exchange,提问作者TheZelucho
相关产品推荐
相关产品推荐

