Scrapy爬虫URL获取问题:代码片段故障排查求助
Hey there, let's work through your Scrapy URL fetching issue together. First, here's your provided code snippet (note it gets cut off mid-line, which is probably part of the problem):
# -*- coding: utf-8 -*- import scrapy from scrapy.spiders import CrawlSpider, Rule from scrapy.linkextractors import LinkExtractor from infoseeker.items import InfoseekerItem as InfoItem class SeekerSpider(CrawlSpider): name = 'seeker' allowed_domains = ['info.mzalendo.com'] start_urls = ['http://info.mzalendo.com/position/member-national-assembly/?page=1'] main_url = 'http://info.mzalendo.com/position/member-national-assembly/' urls = [] retrieving = Fa...
Here are the key areas to check and fix:
1. Complete your truncated code
Your code cuts off at retrieving = Fa... — this line is likely supposed to initialize a URL-fetching or list-building utility (like a custom function or class). Without this logic, your crawler can't properly generate or retrieve additional URLs. First, fill in this missing code to ensure your URL list is being populated correctly.
2. Add CrawlSpider Rules (critical!)
Since you're using CrawlSpider, you haven't defined any Rule objects to tell Scrapy which links to extract and follow. Right now, your crawler will only hit the single URL in start_urls and stop. Add rules to handle pagination or detail links, for example:
rules = ( # Follow pagination links (matches ?page=2, ?page=3, etc.) Rule(LinkExtractor(allow=r'page=\d+'), follow=True, callback='parse_item'), )
Make sure to define the parse_item method to process the response data once you've followed the links.
3. Verify URL generation logic
You've declared urls = [] and main_url, but there's no code to populate the urls list. If you're trying to generate a list of pagination URLs manually, add a loop to build them, then either extend start_urls or override the start_requests method to yield requests for each URL:
def start_requests(self): # Generate pages 1 to 10 (adjust range as needed) for page_num in range(1, 11): url = f"{self.main_url}?page={page_num}" yield scrapy.Request(url, callback=self.parse_item)
4. Check for URL/redirect issues
Your start_urls uses http — confirm if the target site redirects to https. If so, update your URLs to use https to avoid unnecessary redirects that might cause fetch failures.
5. Rule out anti-crawler blocks
Many sites block default Scrapy user agents. Add a realistic user agent in your settings.py:
USER_AGENT = "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
You can also add a small download delay to avoid overwhelming the server:
DOWNLOAD_DELAY = 2
内容的提问来源于stack exchange,提问作者Sam B.

