You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫无法获取目标页面状态与邮箱数据问题求助

Fixing Your Scrapy Spider for Status and Email Extraction

Let's break down what's going wrong with your spider and fix it to correctly grab the Status and email from the lawyer detail pages:

Key Issues in Your Current Code

  • Absolute XPath instead of relative: When you use //span[contains(text(), 'Status:')]//div inside the loop over detail[i], this searches the entire document instead of the current div.line_list_K node. You need to use a relative path starting with ./ to target elements within the current node.
  • Unnecessary loop: Each detail page has exactly one Status entry and one email (if available), so looping through all div.line_list_K elements is redundant and causes incorrect targeting.
  • Incorrect content extraction: Using get() returns the full HTML tag, not just the text inside it. You need to use .//text() to extract the actual text content.
  • Missing email extraction logic: Your code doesn't include any code to scrape the email address.

Corrected Spider Code

import scrapy
from scrapy.http import Request

class TestSpider(scrapy.Spider):
    name = 'test'
    start_urls = ['https://rejestradwokatow.pl/adwokat/list/strona/1/sta/2,3,9']
    custom_settings = {
        'CONCURRENT_REQUESTS_PER_DOMAIN': 1,
        'DOWNLOAD_DELAY': 1,
        'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.130 Safari/537.36'
    }

    def parse(self, response):
        # Extract all lawyer detail page links
        book_links = response.xpath("//td[@class='icon_link']//a/@href").getall()
        for link in book_links:
            full_url = response.urljoin(link)
            yield Request(full_url, callback=self.parse_book)

    def parse_book(self, response):
        # Extract Status text
        status = response.xpath("//span[contains(text(), 'Status:')]/following-sibling::div/text()").get()
        # Extract Email text (handle cases where email might not exist)
        email = response.xpath("//span[contains(text(), 'E-mail:')]/following-sibling::div/a/text()").get()

        # Clean up the extracted text (remove extra whitespace)
        clean_status = status.strip() if status else "No status found"
        clean_email = email.strip() if email else "No email found"

        # Print or yield the results
        print(f"Status: {clean_status}")
        print(f"Email: {clean_email}")

        # Optional: Yield as an item to save to output (like JSON/CSV)
        yield {
            'status': clean_status,
            'email': clean_email,
            'url': response.url
        }

What's Changed:

  1. Simplified link extraction: Used getall() instead of extract() (they're equivalent, but getall() is more readable in modern Scrapy).
  2. Direct targeting of Status and Email: Instead of looping through divs, we directly target the span containing "Status:"/"E-mail:" and grab its sibling div's text. This is more reliable.
  3. Text extraction fix: Used /text() to get only the text content, not the HTML tag.
  4. Error handling: Added checks to handle cases where status or email might be missing (returns a fallback message instead of None).
  5. Optional item yielding: Added code to yield the results as a dictionary, which lets you save the data to files (using scrapy crawl test -o lawyers.json for example).

Testing the Fix

When you run this spider, it will correctly output the Status and Email for each lawyer page. For the example link you provided (https://rejestradwokatow.pl/adwokat/abaewicz-dominik-49965), it should return:

Status: Aktywne
Email: dominik.abaewicz@adwokat.pl

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 01:42:31