Scrapy爬虫未跟进链接:需先校验列表项再爬取详情页
Hey there! Let's work through this Scrapy problem together. You mentioned your base setup (from a tutorial) can scrape the homepage and follow links, but you need to add a check for three required fields (company name, phone number, website) before navigating to detail pages for email extraction. Here's how to adjust your spider to make this work:
Step 1: Add Conditional Validation in the Homepage Parse Method
First, modify your spider's parse function to extract the three fields from each list item, validate they exist, and only yield a request to the detail page if all conditions are met.
import scrapy class CompanySpider(scrapy.Spider): name = "company_spider" start_urls = ["https://your-target-site.com/homepage"] def parse(self, response): # Loop through each company list item (update selector to match your site) for list_item in response.css("div.company-listing-item"): # Extract required fields with fallbacks to avoid None errors company_name = list_item.css("h3.company-name::text").get() phone_number = list_item.css("span.contact-phone::text").get() website = list_item.css("a.company-website::attr(href)").get() # Trim whitespace if fields exist if company_name: company_name = company_name.strip() if phone_number: phone_number = phone_number.strip() # Validate all three fields are present if company_name and phone_number and website: # Get the detail page link (update selector to match your site) detail_link = list_item.css("a.view-detail::attr(href)").get() if detail_link: # Convert relative URL to absolute absolute_detail_url = response.urljoin(detail_link) # Pass the already extracted data to the detail parser via meta yield scrapy.Request( url=absolute_detail_url, callback=self.parse_detail, meta={ "company_data": { "name": company_name, "phone": phone_number, "website": website } } )
Step 2: Parse the Detail Page for Email
Create a parse_detail method to extract the email from the detail page and combine it with the data you already collected:
def parse_detail(self, response): # Retrieve the company data passed from the homepage company_data = response.meta.get("company_data", {}) # Extract email (update selector to match your site's email element) email = response.css("span.company-email::text").get() if email: email = email.strip() # Add email to the data and yield the final item company_data["email"] = email yield company_data
Common Troubleshooting Checks
If your spider still isn't following links, double-check these points:
- Selector Accuracy: Use Scrapy Shell (
scrapy shell "your-homepage-url") to test your CSS/XPath selectors. Make sure they correctly target list items, the three required fields, and the detail page link. - URL Handling: Ensure the detail link is converted to an absolute URL with
response.urljoin()—relative links won't work if you pass them directly toscrapy.Request. - Anti-Circumvention: Some sites block default Scrapy user agents. Add a realistic UA in your
settings.py:USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" - Yield Confirmation: Make sure you're actually yielding the
scrapy.Requestobject inside your conditional check—missing this step means no links get followed.
Once you tweak the selectors to match your target site's HTML structure and implement these checks, your spider should only follow links for qualifying companies and extract the email as expected.
内容的提问来源于stack exchange,提问作者BARNOWL

