You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫URL获取问题:代码片段故障排查求助

Scrapy爬虫URL获取问题排查方案

Hey there, let's work through your Scrapy URL fetching issue together. First, here's your provided code snippet (note it gets cut off mid-line, which is probably part of the problem):

# -*- coding: utf-8 -*-
import scrapy
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from infoseeker.items import InfoseekerItem as InfoItem

class SeekerSpider(CrawlSpider):
    name = 'seeker'
    allowed_domains = ['info.mzalendo.com']
    start_urls = ['http://info.mzalendo.com/position/member-national-assembly/?page=1']
    main_url = 'http://info.mzalendo.com/position/member-national-assembly/'
    urls = []
    retrieving = Fa...

Here are the key areas to check and fix:

1. Complete your truncated code

Your code cuts off at retrieving = Fa... — this line is likely supposed to initialize a URL-fetching or list-building utility (like a custom function or class). Without this logic, your crawler can't properly generate or retrieve additional URLs. First, fill in this missing code to ensure your URL list is being populated correctly.

2. Add CrawlSpider Rules (critical!)

Since you're using CrawlSpider, you haven't defined any Rule objects to tell Scrapy which links to extract and follow. Right now, your crawler will only hit the single URL in start_urls and stop. Add rules to handle pagination or detail links, for example:

rules = (
    # Follow pagination links (matches ?page=2, ?page=3, etc.)
    Rule(LinkExtractor(allow=r'page=\d+'), follow=True, callback='parse_item'),
)

Make sure to define the parse_item method to process the response data once you've followed the links.

3. Verify URL generation logic

You've declared urls = [] and main_url, but there's no code to populate the urls list. If you're trying to generate a list of pagination URLs manually, add a loop to build them, then either extend start_urls or override the start_requests method to yield requests for each URL:

def start_requests(self):
    # Generate pages 1 to 10 (adjust range as needed)
    for page_num in range(1, 11):
        url = f"{self.main_url}?page={page_num}"
        yield scrapy.Request(url, callback=self.parse_item)

4. Check for URL/redirect issues

Your start_urls uses http — confirm if the target site redirects to https. If so, update your URLs to use https to avoid unnecessary redirects that might cause fetch failures.

5. Rule out anti-crawler blocks

Many sites block default Scrapy user agents. Add a realistic user agent in your settings.py:

USER_AGENT = "Mozilla/5.0 (Macintosh; Intel Mac OS X 13_4) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"

You can also add a small download delay to avoid overwhelming the server:

DOWNLOAD_DELAY = 2

内容的提问来源于stack exchange,提问作者Sam B.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:26:31