Scrapy能否抓取通过:hover触发显示的额外内容?爬虫无法提取hover显示邮箱的问题排查
Hey there! Let's work through your two questions step by step—they're great ones for anyone working with dynamic content in Scrapy.
1. Can Scrapy capture HTML content revealed by the :hover pseudo-class?
It depends entirely on how that hover-triggered content is delivered:
- If the content is already in the static HTML (just hidden with CSS)
Scrapy pulls down the full raw HTML of the page, even elements that are hidden with rules likedisplay: noneorvisibility: hidden(triggered by:hover). So you can extract this content directly usingresponse.css(),response.xpath(), or regex onresponse.text—no fancy tricks needed. - If the content loads dynamically via AJAX when you hover
If hovering triggers a background API call to fetch the content (like the email), Scrapy's default setup won't catch this automatically. You'll need to:- Use your browser's DevTools (Network tab) to find the exact AJAX endpoint that returns the data.
- Mimic that request in your spider using
scrapy.Request. - Or, use tools like Selenium or Playwright to simulate the hover action and render the fully interactive page, then extract from the rendered DOM.
2. Why aren't you getting any email results, and how to fix it?
Looking at your code, there are a few key issues stopping you from capturing those emails, plus some context about how biorxiv handles author emails:
Critical Issues in Your Current Code
- Your
Ruleisn't matching any links
YourLinkExtractoris set toallow=r'Items/', but biorxiv has zero URLs containingItems/. That means yourparse_itemmethod is never running—the spider just fetches your start URL and stops, since there's no links to follow that fit the rule. - Variable name typo
You definedemals = re.findall(...)but tried to printprint(emails)—this would throw aNameError(check your spider logs, you'll see it!).
How to Fix the Spider (and Capture Emails)
On biorxiv, author emails are usually either hidden in the static HTML (via CSS) or loaded via a quick AJAX call when you hover. Let's adjust your spider to handle both cases:
Fixed Code (Static HTML First)
Let's switch to a regular Spider instead of CrawlSpider (simpler for your single start URL):
import re import scrapy class MailsSpider(scrapy.Spider): name = 'mails' allowed_domains = ['biorxiv.org'] start_urls = ['https://www.biorxiv.org/content/10.1101/2022.02.28.482253v3'] def parse(self, response): # Check for emails stored in data attributes (common on biorxiv) emails_from_attrs = response.css('[data-email]::attr(data-email)').getall() # Also scan the full page text with regex emails_from_regex = re.findall(r'[\w\.+-]+@[\w\.-]+', response.text) # Combine and remove duplicates all_emails = list(set(emails_from_attrs + emails_from_regex)) self.logger.info(f"Found {len(all_emails)} emails for {response.url}") print(response.url) print(all_emails) # If no emails show up, they're probably loaded via AJAX # Uncomment this section and adjust the API endpoint (find it via DevTools) # author_ids = response.css('[data-author-id]::attr(data-author-id)').getall() # for author_id in author_ids: # yield scrapy.Request( # url=f"https://www.biorxiv.org/api/author/{author_id}/email", # callback=self.parse_author_email # ) # def parse_author_email(self, response): # email = response.json().get('email') # if email: # print(f"Fetched email: {email}")
If Emails Require Hover (Dynamic Rendering)
If the static HTML approach doesn't work, use Selenium to simulate the hover action:
import re from selenium import webdriver from selenium.webdriver.common.action_chains import ActionChains import scrapy class MailsSpider(scrapy.Spider): name = 'mails' allowed_domains = ['biorxiv.org'] start_urls = ['https://www.biorxiv.org/content/10.1101/2022.02.28.482253v3'] def __init__(self): # Make sure ChromeDriver is installed and in your system PATH self.driver = webdriver.Chrome() def closed(self, reason): # Clean up the browser when the spider finishes self.driver.quit() def parse(self, response): self.driver.get(response.url) # Find all author name elements that trigger email on hover author_elements = self.driver.find_elements_by_css_selector('.author-name') emails = [] for elem in author_elements: # Simulate hovering over the author name ActionChains(self.driver).move_to_element(elem).perform() # Grab the email from the popup that appears email_elem = self.driver.find_element_by_css_selector('.email') emails.append(email_elem.text) # Remove duplicates and print results unique_emails = list(set(emails)) print(response.url) print(unique_emails)
内容的提问来源于stack exchange,提问作者Maksim Braelyn

