You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫未跟进链接:需先校验列表项再爬取详情页

Hey there! Let's work through this Scrapy problem together. You mentioned your base setup (from a tutorial) can scrape the homepage and follow links, but you need to add a check for three required fields (company name, phone number, website) before navigating to detail pages for email extraction. Here's how to adjust your spider to make this work:

Step 1: Add Conditional Validation in the Homepage Parse Method

First, modify your spider's parse function to extract the three fields from each list item, validate they exist, and only yield a request to the detail page if all conditions are met.

import scrapy

class CompanySpider(scrapy.Spider):
    name = "company_spider"
    start_urls = ["https://your-target-site.com/homepage"]

    def parse(self, response):
        # Loop through each company list item (update selector to match your site)
        for list_item in response.css("div.company-listing-item"):
            # Extract required fields with fallbacks to avoid None errors
            company_name = list_item.css("h3.company-name::text").get()
            phone_number = list_item.css("span.contact-phone::text").get()
            website = list_item.css("a.company-website::attr(href)").get()

            # Trim whitespace if fields exist
            if company_name:
                company_name = company_name.strip()
            if phone_number:
                phone_number = phone_number.strip()

            # Validate all three fields are present
            if company_name and phone_number and website:
                # Get the detail page link (update selector to match your site)
                detail_link = list_item.css("a.view-detail::attr(href)").get()
                if detail_link:
                    # Convert relative URL to absolute
                    absolute_detail_url = response.urljoin(detail_link)
                    # Pass the already extracted data to the detail parser via meta
                    yield scrapy.Request(
                        url=absolute_detail_url,
                        callback=self.parse_detail,
                        meta={
                            "company_data": {
                                "name": company_name,
                                "phone": phone_number,
                                "website": website
                            }
                        }
                    )

Step 2: Parse the Detail Page for Email

Create a parse_detail method to extract the email from the detail page and combine it with the data you already collected:

def parse_detail(self, response):
        # Retrieve the company data passed from the homepage
        company_data = response.meta.get("company_data", {})
        
        # Extract email (update selector to match your site's email element)
        email = response.css("span.company-email::text").get()
        if email:
            email = email.strip()
        
        # Add email to the data and yield the final item
        company_data["email"] = email
        yield company_data

Common Troubleshooting Checks

If your spider still isn't following links, double-check these points:

  • Selector Accuracy: Use Scrapy Shell (scrapy shell "your-homepage-url") to test your CSS/XPath selectors. Make sure they correctly target list items, the three required fields, and the detail page link.
  • URL Handling: Ensure the detail link is converted to an absolute URL with response.urljoin()—relative links won't work if you pass them directly to scrapy.Request.
  • Anti-Circumvention: Some sites block default Scrapy user agents. Add a realistic UA in your settings.py:
    USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    
  • Yield Confirmation: Make sure you're actually yielding the scrapy.Request object inside your conditional check—missing this step means no links get followed.

Once you tweak the selectors to match your target site's HTML structure and implement these checks, your spider should only follow links for qualifying companies and extract the email as expected.

内容的提问来源于stack exchange,提问作者BARNOWL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:22:56