Scrapy结合登录与商品数据采集功能无法正常运行是什么原因?
问题描述
我尝试从某网站采集商品货号、定价、库存等数据并导出到Excel表格。
如下脚本可成功登录,未登录状态下仅商品货号可见,我此前测试该采集器已可成功抓取商品货号,但在如下示例中将登录和数据采集功能结合后,脚本无法正常运行。
请问我哪里出错了?
import scrapy import pandas as pd from scrapy import FormRequest import os artkl_list = [] price_list = [] stock_list = [] link_site = [] class PostsSpider(scrapy.Spider): name = "posts" start_urls = [ 'https://dealerportal.exertis.nl/action/products/catalog/?producttypeID=5' ] def parseAfterLogin(self, response): # i am not sure if all the syntax below is correct. I can supply the HTML for you to check. for i in response.css('div.productlistblock.row'): artkl = i.css('div.articlenumber::text').extract() price = i.css('span.unitfourprice::text').get().strip() #i am not sure if the syntax here is correct stock = i.css('div.right.stockStatusTitle::text').extract() #i am not sure if the syntax here is correct link = i.css('a.product_img_link::attr(href)').get() #i am not sure if the syntax here is correct # put it names in list artkl_list.append(artkl) price_list.append(price) stock_list.append(stock) link_site.append(link) # display information when you scraped from website print('artkl = ', artkl) print('price = ', price) print('stock = ', stock) print('link_site = ', link) print('\n --------------------------------------------- \n') print('artkl = ', len(artkl_list)) print('price = ', len(price_list)) print('stock = ', len(stock_list)) print('link = ', len(link_site)) print('\n --------------------------------------------- \n') # move to next page next_page = response.css('a.next::attr(href)').get() if next_page: yield response.follow('' + str(next_page)) # put it in dataframe df = pd.DataFrame({ 'artkl': artkl_list, 'price': price_list, 'stock': stock_list, 'link_site': link_site }) # save in excel df.to_excel('exertis.xlsx', index=False) def parse(self, response): os.environ['my_em'] = 'thisismyusername' os.environ['my_pw'] = 'thisismypassword' self.em = os.environ.get('my_em') self.pw = os.environ.get('my_pw') self.login_url = "https://dealerportal.exertis.nl/action/frontusers/login" dataLogin = { 'username': self.em, 'password': self.pw, 'login': 'Inloggen' } print(self.login_url) print('--------') print(dataLogin) print('--------') yield FormRequest(url=self.login_url, formdata=dataLogin, callback=self.parseAfterLogin)
问题原因
- 手动构造登录表单容易遗漏站点隐藏验证字段(比如CSRF令牌),大概率实际登录未成功,没有获取到查看价格、库存的权限
- 登录成功后没有跳转至目标商品列表页:登录表单提交后的回调直接指向商品解析方法,但登录成功后站点默认返回的是登录成功页/首页,不是你要采集的商品列表页,解析方法拿不到对应数据自然不会有输出
- 翻页请求未指定回调函数:翻页调用
response.follow时没有指定回调参数,下一页的响应会默认走parse方法,触发重复登录逻辑,不会解析商品数据 - 字段提取缺少空值校验:
price字段直接对get()返回结果调用strip(),如果匹配不到元素返回None会直接抛出异常导致脚本崩溃;artkl和stock用extract()会返回列表,存入Excel会自带方括号,不符合预期 - 全局变量存数据存在隐患:Scrapy异步请求场景下全局列表容易出现数据混乱,建议换成类属性存储采集结果
- 导出逻辑冗余:当前每一页爬取完成后都会覆盖写入一次Excel,虽然最终结果不受影响,但完全没有必要,建议所有页爬完后再统一导出
修复后代码
import scrapy import pandas as pd from scrapy import FormRequest class PostsSpider(scrapy.Spider): name = "posts" # 目标商品列表初始页 target_catalog_url = 'https://dealerportal.exertis.nl/action/products/catalog/?producttypeID=5' # 类属性存采集结果 artkl_list = [] price_list = [] stock_list = [] link_site = [] def start_requests(self): # 先请求登录页 login_url = "https://dealerportal.exertis.nl/action/frontusers/login" yield scrapy.Request(url=login_url, callback=self.do_login) def do_login(self, response): # 自动提取表单字段提交登录,避免漏传隐藏验证参数 dataLogin = { 'username': '替换为你的实际账号', 'password': '替换为你的实际密码', 'login': 'Inloggen' } yield FormRequest.from_response( response, formdata=dataLogin, callback=self.after_login ) def after_login(self, response): # 登录成功后跳转到商品列表页 yield scrapy.Request(url=self.target_catalog_url, callback=self.parse_catalog) def parse_catalog(self, response): # 解析商品列表 for i in response.css('div.productlistblock.row'): artkl = i.css('div.articlenumber::text').get('').strip() price = i.css('span.unitfourprice::text').get('').strip() stock = i.css('div.right.stockStatusTitle::text').get('').strip() link = i.css('a.product_img_link::attr(href)').get('') self.artkl_list.append(artkl) self.price_list.append(price) self.stock_list.append(stock) self.link_site.append(link) # 打印采集结果 print('artkl = ', artkl) print('price = ', price) print('stock = ', stock) print('link_site = ', link) print('\n --------------------------------------------- \n') print('已采集商品数量:', len(self.artkl_list)) print('\n --------------------------------------------- \n') # 翻页逻辑 next_page = response.css('a.next::attr(href)').get() if next_page: yield response.follow(next_page, callback=self.parse_catalog) else: # 所有页爬完后统一导出Excel df = pd.DataFrame({ 'artkl': self.artkl_list, 'price': self.price_list, 'stock': self.stock_list, 'link_site': self.link_site }) df.to_excel('exertis.xlsx', index=False)
内容的提问来源于stack exchange,提问作者Kyai Fadillah
相关产品推荐
相关产品推荐

