You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy结合登录与商品数据采集功能无法正常运行是什么原因?

问题描述

我尝试从某网站采集商品货号、定价、库存等数据并导出到Excel表格。
如下脚本可成功登录,未登录状态下仅商品货号可见,我此前测试该采集器已可成功抓取商品货号,但在如下示例中将登录和数据采集功能结合后,脚本无法正常运行。
请问我哪里出错了?

import scrapy
import pandas as pd
from scrapy import FormRequest
import os

artkl_list = []
price_list = []
stock_list = []
link_site = []

class PostsSpider(scrapy.Spider):
    name = "posts"

    start_urls = [
        'https://dealerportal.exertis.nl/action/products/catalog/?producttypeID=5'
       ]

    def parseAfterLogin(self, response):
        # i am not sure if all the syntax below is correct. I can supply the HTML for you to check.
        for i in response.css('div.productlistblock.row'):
            artkl = i.css('div.articlenumber::text').extract()
            price = i.css('span.unitfourprice::text').get().strip()   #i am not sure if the syntax here is correct
            stock = i.css('div.right.stockStatusTitle::text').extract()   #i am not sure if the syntax here is correct
            link = i.css('a.product_img_link::attr(href)').get()   #i am not sure if the syntax here is correct

            # put it names in list
            artkl_list.append(artkl)
            price_list.append(price)
            stock_list.append(stock)
            link_site.append(link)

            # display information when you scraped from website
            print('artkl              = ', artkl)
            print('price              = ', price)
            print('stock              = ', stock)
            print('link_site          = ', link)

            print('\n --------------------------------------------- \n')
            print('artkl             = ', len(artkl_list))
            print('price             = ', len(price_list))
            print('stock             = ', len(stock_list))
            print('link              = ', len(link_site))
            print('\n --------------------------------------------- \n')

        # move to next page
        next_page = response.css('a.next::attr(href)').get()
        if next_page:
            yield response.follow('' + str(next_page))

        # put it in dataframe
        df = pd.DataFrame({
            'artkl': artkl_list,
            'price': price_list,
            'stock': stock_list,
            'link_site': link_site
        })
        # save in excel
        df.to_excel('exertis.xlsx', index=False)

    def parse(self, response):
        os.environ['my_em'] = 'thisismyusername'
        os.environ['my_pw'] = 'thisismypassword'
        self.em = os.environ.get('my_em')
        self.pw = os.environ.get('my_pw')

        self.login_url = "https://dealerportal.exertis.nl/action/frontusers/login"

        dataLogin = {
            'username': self.em,
            'password': self.pw,
            'login': 'Inloggen'
        }
        print(self.login_url)
        print('--------')
        print(dataLogin)
        print('--------')
        yield FormRequest(url=self.login_url, formdata=dataLogin, callback=self.parseAfterLogin)
问题原因
  • 手动构造登录表单容易遗漏站点隐藏验证字段(比如CSRF令牌),大概率实际登录未成功,没有获取到查看价格、库存的权限
  • 登录成功后没有跳转至目标商品列表页:登录表单提交后的回调直接指向商品解析方法,但登录成功后站点默认返回的是登录成功页/首页,不是你要采集的商品列表页,解析方法拿不到对应数据自然不会有输出
  • 翻页请求未指定回调函数:翻页调用response.follow时没有指定回调参数,下一页的响应会默认走parse方法,触发重复登录逻辑,不会解析商品数据
  • 字段提取缺少空值校验:price字段直接对get()返回结果调用strip(),如果匹配不到元素返回None会直接抛出异常导致脚本崩溃;artkl和stock用extract()会返回列表,存入Excel会自带方括号,不符合预期
  • 全局变量存数据存在隐患:Scrapy异步请求场景下全局列表容易出现数据混乱,建议换成类属性存储采集结果
  • 导出逻辑冗余:当前每一页爬取完成后都会覆盖写入一次Excel,虽然最终结果不受影响,但完全没有必要,建议所有页爬完后再统一导出
修复后代码
import scrapy
import pandas as pd
from scrapy import FormRequest

class PostsSpider(scrapy.Spider):
    name = "posts"
    # 目标商品列表初始页
    target_catalog_url = 'https://dealerportal.exertis.nl/action/products/catalog/?producttypeID=5'
    # 类属性存采集结果
    artkl_list = []
    price_list = []
    stock_list = []
    link_site = []

    def start_requests(self):
        # 先请求登录页
        login_url = "https://dealerportal.exertis.nl/action/frontusers/login"
        yield scrapy.Request(url=login_url, callback=self.do_login)

    def do_login(self, response):
        # 自动提取表单字段提交登录,避免漏传隐藏验证参数
        dataLogin = {
            'username': '替换为你的实际账号',
            'password': '替换为你的实际密码',
            'login': 'Inloggen'
        }
        yield FormRequest.from_response(
            response,
            formdata=dataLogin,
            callback=self.after_login
        )

    def after_login(self, response):
        # 登录成功后跳转到商品列表页
        yield scrapy.Request(url=self.target_catalog_url, callback=self.parse_catalog)

    def parse_catalog(self, response):
        # 解析商品列表
        for i in response.css('div.productlistblock.row'):
            artkl = i.css('div.articlenumber::text').get('').strip()
            price = i.css('span.unitfourprice::text').get('').strip()
            stock = i.css('div.right.stockStatusTitle::text').get('').strip()
            link = i.css('a.product_img_link::attr(href)').get('')

            self.artkl_list.append(artkl)
            self.price_list.append(price)
            self.stock_list.append(stock)
            self.link_site.append(link)

            # 打印采集结果
            print('artkl              = ', artkl)
            print('price              = ', price)
            print('stock              = ', stock)
            print('link_site          = ', link)
            print('\n --------------------------------------------- \n')
            print('已采集商品数量:', len(self.artkl_list))
            print('\n --------------------------------------------- \n')

        # 翻页逻辑
        next_page = response.css('a.next::attr(href)').get()
        if next_page:
            yield response.follow(next_page, callback=self.parse_catalog)
        else:
            # 所有页爬完后统一导出Excel
            df = pd.DataFrame({
                'artkl': self.artkl_list,
                'price': self.price_list,
                'stock': self.stock_list,
                'link_site': self.link_site
            })
            df.to_excel('exertis.xlsx', index=False)

内容的提问来源于stack exchange,提问作者Kyai Fadillah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 03:09:02