Scrapy爬虫开发:如何跳过指定<a>标签并选中目标链接
解决Scrapy分页中定位正确标签的问题
针对你遇到的分页链接定位错误问题,结合提供的HTML结构,给出以下几种可靠的解决方案:
方案1:通过位置定位下一页链接
观察HTML结构,下一页的<a>标签是nav节点下的最后一个同级<a>元素,直接用last()函数定位:
next_page_href = response.xpath("//nav[@class='mp-PaginationControls-pagination']/a[last()]/@href").get()
如果确认它始终是第二个<a>标签,也可以用索引定位:
next_page_href = response.xpath("//nav[@class='mp-PaginationControls-pagination']/a[2]/@href").get()
方案2:通过子元素class精准筛选(更稳定)
下一页按钮的<a>标签内部包含带右箭头图标的<span>,其class为mp-svg-arrow-right--inverse,通过这个子元素特征定位父级<a>:
next_page_href = response.xpath("//nav[@class='mp-PaginationControls-pagination']/a[span[@class='mp-Button-icon mp-Button-icon--center mp-svg-arrow-right--inverse']]/@href").get()
如果担心class名有变动,也可以用contains匹配部分class:
next_page_href = response.xpath("//nav[@class='mp-PaginationControls-pagination']/a[span[contains(@class, 'mp-svg-arrow-right--inverse')]]/@href").get()
额外注意
获取到相对路径后,记得用Scrapy的response.urljoin()转换为绝对URL,避免跳转异常:
if next_page_href: next_page_url = response.urljoin(next_page_href) yield scrapy.Request(next_page_url, callback=self.parse)
原XPath//nav[@class='mp-PaginationControls-pagination']/a/@href只取第一个匹配的<a>(即左箭头的返回按钮),导致进入第二页后继续点击返回按钮,引发循环或异常,以上方案可精准定位到下一页的蓝色高亮按钮。
内容的提问来源于stack exchange,提问作者Sean
相关产品推荐
相关产品推荐

