You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Scrapy XML Feed Spider仅提取站点地图中的商品页URL

调整Scrapy XML Feed Spider仅提取商品页URL

你遇到的问题是原代码使用了全局的//loc/text() XPath表达式,会匹配所有<loc>标签(包括<image:loc>移除命名空间后的标签),导致同时抓取到商品页和图片URL。以下是两种调整方案:

方案1:移除命名空间后使用相对路径

直接在当前<url>节点范围内查找直接子级的<loc>标签,避免全局搜索:

start_urls = ['https://www.example.com/sitemap.xml']
namespaces = [('n', 'http://www.sitemaps.org/schemas/sitemap/0.9')]
itertag = 'url'

def parse_node(self, response, selector):
    selector.remove_namespaces()
    # 仅获取当前<url>节点下的直接子级<loc>文本
    product_url = selector.xpath('./loc/text()').get()
    if product_url:
        yield {'url': product_url}

方案2:保留命名空间使用限定路径

如果不想移除命名空间,可以通过指定命名空间前缀精准匹配目标标签:

start_urls = ['https://www.example.com/sitemap.xml']
namespaces = [('n', 'http://www.sitemaps.org/schemas/sitemap/0.9')]
itertag = 'url'

def parse_node(self, response, selector):
    # 利用命名空间前缀匹配<url>节点下的<n:loc>标签
    product_url = selector.xpath('./n:loc/text()', namespaces=self.namespaces).get()
    if product_url:
        yield {'url': product_url}

关键说明

  • 原代码里的//loc/text()是全局搜索,会遍历整个XML文档的所有<loc>标签,包括图片站点地图的<image:loc>(移除命名空间后会被识别为<loc>)。
  • 使用./loc(或./n:loc)是相对路径,仅在当前处理的<url>节点范围内查找直接子级的<loc>,这正是商品页URL的所在位置。
  • 添加if product_url:判断是为了过滤掉可能没有<loc>标签的无效<url>节点,避免输出空值。

内容的提问来源于stack exchange,提问作者user3125823

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 09:40:31