You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫中print_urls()函数无法执行问题求助

问题原因及解决办法

核心问题:生成器函数未被迭代

你的print_urls()是生成器函数(内部使用了yield),直接调用self.print_urls()不会触发函数内的代码执行——生成器只有在被迭代(比如for循环、yield from语法)时才会运行。

具体修改步骤

1. 修复handle_inscriptions中的调用方式

把直接调用生成器的代码,改成用yield from迭代生成器,这样既能触发print_urls执行,又能把生成的Item返回给Scrapy:

def handle_inscriptions(self, response):
    homes, success = self.success(response)
    if success == True:
        print(Fore.GREEN + 'Count ' + str(homes['d']['Result']['count']) + Style.RESET_ALL)
    # 直接创建Selector并传入函数,不用实例变量存储
    html_selector = Selector(text=homes['d']['Result']['html'])
    # 迭代生成器并返回结果
    yield from self.print_urls(html_selector)

2. 修改print_urls函数,用参数传递Selector

不要用实例变量self.html传递数据,改成参数传入更安全,避免多请求并发时的变量覆盖问题:

def print_urls(self, page_html):
    print('try')
    homes = page_html.xpath('//div[contains(@class, "property-thumbnail-item")]')
    for home in homes:
        yield {
            'home_url': home.xpath('.//a[@class="property-thumbnail-summary-link"]/@href').get()
        }

3. 修复success函数的返回值漏洞

当请求失败时,原函数只返回False,但调用时用了homes, success = self.success(response)的解包语法,会触发报错。修改为统一返回两个值:

def success(self, response):
    my_dict = literal_eval(response.body.decode('utf-8').replace(':true}', ':True}'))
    if my_dict['d']['Succeeded'] == True:
        return my_dict, True
    else:
        return None, False  # 统一返回两个值,避免解包错误

额外提示

  • Scrapy的Spider回调函数通过yield返回Item或Request,生成器函数的结果必须被迭代才能生效。
  • 尽量避免用实例变量存储请求相关数据,多请求并发时容易出现数据污染。

内容的提问来源于stack exchange,提问作者Carter James

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 06:35:48