如何在Scrapy中通过Rule对象仅提取div元素的data-next属性链接(排除子a标签链接)
Great question! I've run into this exact scenario before—Scrapy's LinkExtractor has some default behaviors that can trip you up here, but there's a clean way to make it work exactly how you want.
Why Your Original Approach Didn't Work
Your first attempt using restrict_xpaths=('//div[@data-next]/@data-next') failed because LinkExtractor expects restrict_xpaths to point to elements (not direct attribute values) where it should look for links. When you removed the @data-next, it defaulted to scanning the <a> tags inside the div (since LinkExtractor targets <a> and <area> tags by default), hence grabbing all those extra links.
The Fix: Customize LinkExtractor's Tags and Attributes
To tell LinkExtractor to ignore the child <a> tags and only pull the data-next attribute from the parent div, you need to adjust two key parameters:
tags: Override the default to target<div>instead of<a>/<area>attrs: Specify that you want to extract thedata-nextattribute instead of the defaulthref
Here's the working Rule code:
Rule(LinkExtractor( tags=['div'], attrs=['data-next'], restrict_xpaths='//div[@data-next]' ), callback='parse_item')
How This Works
tags=['div']: TellsLinkExtractorto only consider<div>elements when looking for linksattrs=['data-next']: Instructs it to extract the value of thedata-nextattribute from those divsrestrict_xpaths='//div[@data-next]': Narrows it down to only divs that actually have adata-nextattribute (so you don't waste time scanning other divs)
With this setup, Scrapy will only grab link0 from the div's data-next attribute and completely ignore the child <a> tags' href values.
内容的提问来源于stack exchange,提问作者Jonathan Simpson

