You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Scrapy爬取详情页而非列表页?附示例代码

如何用Scrapy爬取目标网站详情页内容

现有代码的问题

  • allowed_domains配置错误,目标域名是ssr1.scrape.center,原配置的域名会导致合法请求被Scrapy过滤
  • 详情页解析的XPath存在语法错误(如address字段的XPath开头缺失//)
  • 提取详情链接时使用extract()[0]存在索引越界风险,建议用更安全的get()或extract_first()
  • 下一页链接拼接逻辑有漏洞,当没有下一页时会生成无效URL

修正后的完整代码

import scrapy
from scrapytutorial.items import ScrapytutorialItem


class FirstprojectSpider(scrapy.Spider):
    name = 'firstproject'  
    # 修正允许的域名
    allowed_domains = ['ssr1.scrape.center'] 
    # 替换为目标网站首页,或直接删除(因重写了start_requests)
    start_urls = ['https://ssr1.scrape.center']  

    def start_requests(self):
        # 遍历1-10页列表页
        for page in range(1, 11):
            url = f'https://ssr1.scrape.center/page/{page}'
            yield scrapy.Request(
                url=url,
                headers={
                    'user_agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/105.0.0.0 Safari/537.36',
                    'host': 'ssr1.scrape.center',
                    'Referer': 'https://ssr1.scrape.center/'
                },
                callback=self.parse
            )

    
    def parse(self, response):
        """处理列表页,提取详情页链接"""
        # 优化XPath,通过class定位电影卡片更稳定
        movie_list = response.xpath('//div[contains(@class, "el-card")]')
        for card in movie_list:
            # 安全提取详情链接,避免索引越界
            detail_path = card.xpath('.//a[@class="name"]/@href').get()
            if detail_path:
                # 使用response.follow自动拼接完整URL,代码更简洁
                yield response.follow(
                    url=detail_path,
                    callback=self.getdetail
                )
        # 修正下一页链接提取逻辑
        next_page_path = response.xpath('//ul[@class="el-pager"]/li[@class="btn-next"]/a/@href').get()
        if next_page_path and next_page_path != '#':
            yield response.follow(
                url=next_page_path,
                callback=self.parse
            )

    
    def getdetail(self, response):
        """解析详情页内容"""
        items = ScrapytutorialItem()

        # 优化XPath,通过class定位元素,同时处理空值情况
        items['name'] = response.xpath('//h2[@class="m-b-sm"]/text()').get().strip() if response.xpath('//h2[@class="m-b-sm"]/text()').get() else None
        items['address'] = response.xpath('//div[@class="m-v-sm info"]/span[1]/text()').get().strip() if response.xpath('//div[@class="m-v-sm info"]/span[1]/text()').get() else None
        items['times'] = response.xpath('//div[@class="m-v-sm info"]/span[3]/text()').get().strip() if response.xpath('//div[@class="m-v-sm info"]/span[3]/text()').get() else None
        items['timestamp'] = response.xpath('//div[@class="m-v-sm info"]/span[last()]/text()').get().strip() if response.xpath('//div[@class="m-v-sm info"]/span[last()]/text()').get() else None
        items['plot'] = response.xpath('//div[@class="m-v-sm synopsis"]/p/text()').get().strip() if response.xpath('//div[@class="m-v-sm synopsis"]/p/text()').get() else None
        
        yield items

关键优化点说明

  • 域名配置:allowed_domains必须匹配目标网站域名,否则Scrapy会拦截跨域请求
  • 链接提取:用get()替代extract()[0],元素不存在时返回None,避免抛出索引错误
  • XPath稳定性:优先通过class属性定位元素,比依赖固定id更能适配页面结构的微调
  • URL拼接:使用response.follow自动处理基础URL拼接,无需手动拼接完整链接,代码更简洁

内容的提问来源于stack exchange,提问作者Aloha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 20:35:24