You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy能否抓取通过:hover触发显示的额外内容?爬虫无法提取hover显示邮箱的问题排查

Answers to Your Scrapy Hover & Email Scraping Questions

Hey there! Let's work through your two questions step by step—they're great ones for anyone working with dynamic content in Scrapy.

1. Can Scrapy capture HTML content revealed by the :hover pseudo-class?

It depends entirely on how that hover-triggered content is delivered:

  • If the content is already in the static HTML (just hidden with CSS)
    Scrapy pulls down the full raw HTML of the page, even elements that are hidden with rules like display: none or visibility: hidden (triggered by :hover). So you can extract this content directly using response.css(), response.xpath(), or regex on response.text—no fancy tricks needed.
  • If the content loads dynamically via AJAX when you hover
    If hovering triggers a background API call to fetch the content (like the email), Scrapy's default setup won't catch this automatically. You'll need to:
    1. Use your browser's DevTools (Network tab) to find the exact AJAX endpoint that returns the data.
    2. Mimic that request in your spider using scrapy.Request.
    3. Or, use tools like Selenium or Playwright to simulate the hover action and render the fully interactive page, then extract from the rendered DOM.

2. Why aren't you getting any email results, and how to fix it?

Looking at your code, there are a few key issues stopping you from capturing those emails, plus some context about how biorxiv handles author emails:

Critical Issues in Your Current Code

  1. Your Rule isn't matching any links
    Your LinkExtractor is set to allow=r'Items/', but biorxiv has zero URLs containing Items/. That means your parse_item method is never running—the spider just fetches your start URL and stops, since there's no links to follow that fit the rule.
  2. Variable name typo
    You defined emals = re.findall(...) but tried to print print(emails)—this would throw a NameError (check your spider logs, you'll see it!).

How to Fix the Spider (and Capture Emails)

On biorxiv, author emails are usually either hidden in the static HTML (via CSS) or loaded via a quick AJAX call when you hover. Let's adjust your spider to handle both cases:

Fixed Code (Static HTML First)

Let's switch to a regular Spider instead of CrawlSpider (simpler for your single start URL):

import re
import scrapy

class MailsSpider(scrapy.Spider):
    name = 'mails'
    allowed_domains = ['biorxiv.org']
    start_urls = ['https://www.biorxiv.org/content/10.1101/2022.02.28.482253v3']

    def parse(self, response):
        # Check for emails stored in data attributes (common on biorxiv)
        emails_from_attrs = response.css('[data-email]::attr(data-email)').getall()
        # Also scan the full page text with regex
        emails_from_regex = re.findall(r'[\w\.+-]+@[\w\.-]+', response.text)
        
        # Combine and remove duplicates
        all_emails = list(set(emails_from_attrs + emails_from_regex))
        
        self.logger.info(f"Found {len(all_emails)} emails for {response.url}")
        print(response.url)
        print(all_emails)

        # If no emails show up, they're probably loaded via AJAX
        # Uncomment this section and adjust the API endpoint (find it via DevTools)
        # author_ids = response.css('[data-author-id]::attr(data-author-id)').getall()
        # for author_id in author_ids:
        #     yield scrapy.Request(
        #         url=f"https://www.biorxiv.org/api/author/{author_id}/email",
        #         callback=self.parse_author_email
        #     )
    
    # def parse_author_email(self, response):
    #     email = response.json().get('email')
    #     if email:
    #         print(f"Fetched email: {email}")

If Emails Require Hover (Dynamic Rendering)

If the static HTML approach doesn't work, use Selenium to simulate the hover action:

import re
from selenium import webdriver
from selenium.webdriver.common.action_chains import ActionChains
import scrapy

class MailsSpider(scrapy.Spider):
    name = 'mails'
    allowed_domains = ['biorxiv.org']
    start_urls = ['https://www.biorxiv.org/content/10.1101/2022.02.28.482253v3']

    def __init__(self):
        # Make sure ChromeDriver is installed and in your system PATH
        self.driver = webdriver.Chrome()

    def closed(self, reason):
        # Clean up the browser when the spider finishes
        self.driver.quit()

    def parse(self, response):
        self.driver.get(response.url)
        # Find all author name elements that trigger email on hover
        author_elements = self.driver.find_elements_by_css_selector('.author-name')
        
        emails = []
        for elem in author_elements:
            # Simulate hovering over the author name
            ActionChains(self.driver).move_to_element(elem).perform()
            # Grab the email from the popup that appears
            email_elem = self.driver.find_element_by_css_selector('.email')
            emails.append(email_elem.text)
        
        # Remove duplicates and print results
        unique_emails = list(set(emails))
        print(response.url)
        print(unique_emails)

内容的提问来源于stack exchange,提问作者Maksim Braelyn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 16:57:44