Scrapy爬虫无法获取目标页面状态与邮箱数据问题求助
Fixing Your Scrapy Spider for Status and Email Extraction
Let's break down what's going wrong with your spider and fix it to correctly grab the Status and email from the lawyer detail pages:
Key Issues in Your Current Code
- Absolute XPath instead of relative: When you use
//span[contains(text(), 'Status:')]//divinside the loop overdetail[i], this searches the entire document instead of the currentdiv.line_list_Knode. You need to use a relative path starting with./to target elements within the current node. - Unnecessary loop: Each detail page has exactly one Status entry and one email (if available), so looping through all
div.line_list_Kelements is redundant and causes incorrect targeting. - Incorrect content extraction: Using
get()returns the full HTML tag, not just the text inside it. You need to use.//text()to extract the actual text content. - Missing email extraction logic: Your code doesn't include any code to scrape the email address.
Corrected Spider Code
import scrapy from scrapy.http import Request class TestSpider(scrapy.Spider): name = 'test' start_urls = ['https://rejestradwokatow.pl/adwokat/list/strona/1/sta/2,3,9'] custom_settings = { 'CONCURRENT_REQUESTS_PER_DOMAIN': 1, 'DOWNLOAD_DELAY': 1, 'USER_AGENT': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/79.0.3945.130 Safari/537.36' } def parse(self, response): # Extract all lawyer detail page links book_links = response.xpath("//td[@class='icon_link']//a/@href").getall() for link in book_links: full_url = response.urljoin(link) yield Request(full_url, callback=self.parse_book) def parse_book(self, response): # Extract Status text status = response.xpath("//span[contains(text(), 'Status:')]/following-sibling::div/text()").get() # Extract Email text (handle cases where email might not exist) email = response.xpath("//span[contains(text(), 'E-mail:')]/following-sibling::div/a/text()").get() # Clean up the extracted text (remove extra whitespace) clean_status = status.strip() if status else "No status found" clean_email = email.strip() if email else "No email found" # Print or yield the results print(f"Status: {clean_status}") print(f"Email: {clean_email}") # Optional: Yield as an item to save to output (like JSON/CSV) yield { 'status': clean_status, 'email': clean_email, 'url': response.url }
What's Changed:
- Simplified link extraction: Used
getall()instead ofextract()(they're equivalent, butgetall()is more readable in modern Scrapy). - Direct targeting of Status and Email: Instead of looping through divs, we directly target the span containing "Status:"/"E-mail:" and grab its sibling div's text. This is more reliable.
- Text extraction fix: Used
/text()to get only the text content, not the HTML tag. - Error handling: Added checks to handle cases where status or email might be missing (returns a fallback message instead of
None). - Optional item yielding: Added code to yield the results as a dictionary, which lets you save the data to files (using
scrapy crawl test -o lawyers.jsonfor example).
Testing the Fix
When you run this spider, it will correctly output the Status and Email for each lawyer page. For the example link you provided (https://rejestradwokatow.pl/adwokat/abaewicz-dominik-49965), it should return:
Status: Aktywne Email: dominik.abaewicz@adwokat.pl
内容的提问来源于stack exchange,提问作者Amen Aziz
相关产品推荐
相关产品推荐

