You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy CrawlSpider处理input表单下一页及提取图片尺寸问题咨询

Scrapy CrawlSpider 表单翻页与尺寸提取解决方案

核心修改说明

1. 表单下一页跳转问题修复

原代码使用LinkExtractor处理表单input标签的逻辑不符合组件设计规则,LinkExtractor原生仅适配a标签等带链接属性的静态资源,无法直接解析表单提交路径。
修改方案:

  • 新增parse_start_url方法统一处理所有列表页响应,手动提取下一页表单内的路径值,拼接为完整URL后发起翻页请求
  • 保留原有Rule规则处理详情页链接提取,不改动原有详情页解析逻辑

2. cm单位尺寸提取修复

原XPath路径定位错误,仅提取了表头文本,未获取到实际尺寸值。
修改方案:

  • 修正XPath路径,定位到Dimensions表头对应的实际值单元格
  • 增加正则匹配过滤,仅保留带cm单位的有效尺寸结果

完整可运行代码

import scrapy
import re
from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule


class ToscrapeSpider(CrawlSpider):
    name = 'toscrape'
    allowed_domains = ['pstrial-2019-12-16.toscrape.com']
    start_urls = ['http://pstrial-2019-12-16.toscrape.com/browse/insunsh']

    # 仅保留详情页提取Rule,翻页逻辑改为手动处理
    rules = (
        Rule(LinkExtractor(restrict_xpaths="//div[@id='body']/div[2]/a"), callback='parse_item', follow=False),
    )

    # 处理列表页(包括初始页和后续翻页)
    def parse_start_url(self, response):
        # 先交给父类处理,触发Rule提取当前页详情页链接
        yield from super().parse_start_url(response)
        # 提取下一页路径
        next_page_path = response.xpath("//form[@class='nav next']/input[1]/@value").get()
        if next_page_path:
            # 自动拼接域名发起下一页请求,回调到当前方法继续处理下一页列表
            yield response.follow(next_page_path, callback=self.parse_start_url)

    def parse_item(self, response):
        # 提取尺寸并过滤cm单位
        dim_raw = response.xpath("//tr/td[text()='Dimensions']/following-sibling::td[1]/text()").get()
        dim_cm = None
        if dim_raw:
            cm_match = re.search(r'(\d+\s*x\s*\d+\s*cm)', dim_raw.strip(), re.I)
            if cm_match:
                dim_cm = cm_match.group(1).strip()
        yield {
            'Image': response.xpath("//div[@id='body']/img/@src").get(),
            'Title': response.xpath("//div[@id='content']/h1/text()").get().strip(),
            'artist': response.xpath("//div[@id='content']/h2/text()").get().strip(),
            'Description': ''.join(response.xpath("//div[@class='description']/p/text()").getall()).strip(),
            'URL': response.url,
            'Dimension': dim_cm
        }

内容的提问来源于stack exchange,提问作者Deepak Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 14:36:03