Scrapy自定义方法未调用问题求助:拆分代码后无法执行
Hey! I see exactly what's going on here—let's break down why your parse_paragraph method isn't running as expected, and how to fix it while keeping your code clean and modular.
Your parse_paragraph is a generator function (it uses yield), but right now you're just calling it like a regular function in parse:
self.parse_paragraph(div_list)
Generators in Python are lazy—they don't execute any code until you iterate over them. Since you're not letting Scrapy (or any loop) iterate over the generator returned by parse_paragraph, none of the code inside that method runs. That's why you don't see your print statement, no items are yielded, and breakpoints don't trigger.
Scrapy only processes iterables returned from callback methods like parse—so you need to pass the generator from parse_paragraph to Scrapy properly.
There are two straightforward ways to fix this without cramming all logic into parse:
Option 1: Use yield from (Python 3.3+)
This is the cleanest approach. yield from tells Python to iterate over the generator returned by parse_paragraph and forward all its yielded items to Scrapy.
Modify your parse method like this:
def parse(self, response): div_list = response.xpath("//div[@class='contentF']/div[@class='content_l']/p") # Forward all items from parse_paragraph to Scrapy yield from self.parse_paragraph(div_list)
Your existing parse_paragraph method doesn't need any changes—keep it exactly as written.
Option 2: Return a List of Items (For Older Python Versions)
If you're stuck on a Python version before 3.3 (unlikely these days, but just in case), adjust parse_paragraph to collect items into a list and return it, then iterate over that list in parse:
def parse(self, response): div_list = response.xpath("//div[@class='contentF']/div[@class='content_l']/p") items = self.parse_paragraph(div_list) for item in items: yield item def parse_paragraph(self, div_list): items = [] for div in div_list: duilian_text_list = div.xpath('./text()').extract() for duilian_text in duilian_text_list: duilian_item = DuilianItem() duilian_item['category_id'] = 1 duilian = duilian_text duilian_item['name'] = duilian duilian_item['desc'] = '' print('I reach here...') items.append(duilian_item) return items
parse Worked When you moved all code into parse, you were yielding items directly from the parse method itself. Scrapy automatically iterates over the generator returned by parse, so all your code ran and items were processed as expected.
内容的提问来源于stack exchange,提问作者zdd

