You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy自定义方法未调用问题求助:拆分代码后无法执行

Hey! I see exactly what's going on here—let's break down why your parse_paragraph method isn't running as expected, and how to fix it while keeping your code clean and modular.

Core Issue: Generators Need to Be Iterated

Your parse_paragraph is a generator function (it uses yield), but right now you're just calling it like a regular function in parse:

self.parse_paragraph(div_list)

Generators in Python are lazy—they don't execute any code until you iterate over them. Since you're not letting Scrapy (or any loop) iterate over the generator returned by parse_paragraph, none of the code inside that method runs. That's why you don't see your print statement, no items are yielded, and breakpoints don't trigger.

Scrapy only processes iterables returned from callback methods like parse—so you need to pass the generator from parse_paragraph to Scrapy properly.

Fixes to Keep Your Code Modular

There are two straightforward ways to fix this without cramming all logic into parse:

Option 1: Use yield from (Python 3.3+)

This is the cleanest approach. yield from tells Python to iterate over the generator returned by parse_paragraph and forward all its yielded items to Scrapy.

Modify your parse method like this:

def parse(self, response):
    div_list = response.xpath("//div[@class='contentF']/div[@class='content_l']/p")
    # Forward all items from parse_paragraph to Scrapy
    yield from self.parse_paragraph(div_list)

Your existing parse_paragraph method doesn't need any changes—keep it exactly as written.

Option 2: Return a List of Items (For Older Python Versions)

If you're stuck on a Python version before 3.3 (unlikely these days, but just in case), adjust parse_paragraph to collect items into a list and return it, then iterate over that list in parse:

def parse(self, response):
    div_list = response.xpath("//div[@class='contentF']/div[@class='content_l']/p")
    items = self.parse_paragraph(div_list)
    for item in items:
        yield item

def parse_paragraph(self, div_list):
    items = []
    for div in div_list:
        duilian_text_list = div.xpath('./text()').extract()
        for duilian_text in duilian_text_list:
            duilian_item = DuilianItem()
            duilian_item['category_id'] = 1
            duilian = duilian_text
            duilian_item['name'] = duilian
            duilian_item['desc'] = ''
            print('I reach here...')
            items.append(duilian_item)
    return items
Why Merging into parse Worked

When you moved all code into parse, you were yielding items directly from the parse method itself. Scrapy automatically iterates over the generator returned by parse, so all your code ran and items were processed as expected.

内容的提问来源于stack exchange,提问作者zdd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:36:04