Scrapy新手求助:针对不同URL调用不同处理函数
Hey there! Nice work getting those target links put together—you’re already off to a solid start with your Scrapy project. Now let’s tackle the tricky part: handling those inconsistent page structures. Here’s what I’d recommend based on my experience:
1. Chain Selectors for Fallback Matching
Even if pages look totally different, they often share core data points (like titles or descriptions). Instead of writing separate parsers right away, use fallback selectors to cover multiple cases. For example:
# In your target page parser title = response.css('h1.product-title::text').get() or response.css('div.item-heading::text').get() or "No title found"
This way, if the first selector doesn’t pick up a value, it tries the next one. Super simple and effective for minor structure differences.
2. Use Item Loaders for Clean, Flexible Data Collection
Item Loaders are a game-changer for messy data structures. They let you define multiple ways to grab the same field, and handle processing automatically. First, set up your Item class:
import scrapy from scrapy.loader import ItemLoader from itemloaders.processors import TakeFirst, Join class MyScraperItem(scrapy.Item): link_name = scrapy.Field(output_processor=TakeFirst()) target_url = scrapy.Field(output_processor=TakeFirst()) content = scrapy.Field(output_processor=Join())
Then, in your spider’s target page callback, use the loader to pull data from multiple selectors:
def parse_target_page(self, response): loader = ItemLoader(item=MyScraperItem(), response=response) # Add all possible selectors for content loader.add_css('content', 'div.main-body p::text') loader.add_css('content', 'article.post-content ::text') loader.add_xpath('content', '//div[@class="product-desc"]//text()') # Pull in the link name you passed via meta loader.add_value('link_name', response.meta.get('link_name')) loader.add_value('target_url', response.url) yield loader.load_item()
Don’t forget to pass the link name when you generate the target request:
def parse(self, response): # Your existing code to extract link_name and link_address for entry in response.css('a.link-item'): link_name = entry.css('::text').get().strip() link_address = entry.css('::attr(href)').get() target_url = response.urljoin(link_address) yield scrapy.Request( target_url, callback=self.parse_target_page, meta={'link_name': link_name} )
3. Route Requests to Specialized Parsers for Major Differences
If you’re dealing with completely distinct page types (like blog posts vs. product pages), add a check in your main target parser to send the response to the right handler:
def parse_target_page(self, response): # Check for product page marker if response.css('div.product-specs').exists(): yield self.parse_product_page(response) # Check for blog page marker elif response.css('div.blog-meta').exists(): yield self.parse_blog_page(response) # Log unknown pages to debug later else: self.logger.warning(f"Unrecognized page structure at {response.url}") def parse_product_page(self, response): # Product-specific parsing logic item = MyScraperItem() item['link_name'] = response.meta.get('link_name') item['target_url'] = response.url item['content'] = ' '.join(response.css('div.product-details ::text').getall()).strip() yield item def parse_blog_page(self, response): # Blog-specific parsing logic item = MyScraperItem() item['link_name'] = response.meta.get('link_name') item['target_url'] = response.url item['content'] = ' '.join(response.css('article.post-body p::text').getall()).strip() yield item
4. Test Selectors Quickly with Scrapy Shell
Don’t guess and check in your code—use the Scrapy shell to test selectors on live pages. Just run:
scrapy shell https://your-target-page-url.com
Then you can experiment with response.css() or response.xpath() commands in real time to find the right selectors before adding them to your spider.
Start small—focus on 2-3 page types first, get those working smoothly, then expand to others. You’ve got this!
内容的提问来源于stack exchange,提问作者srk

