Scrapy爬虫部署到Apify后无法执行aaa_products函数求助
Scrapy爬虫本地正常但部署到Apify后无法触发aaa_products函数的问题
问题现象
- 本地运行正常:爬虫可正常解析商品页面,输出预期结果,日志示例:
2023-10-08 00:15:41 [scrapy.core.scraper] DEBUG: Scraped from <200 https://www.apc.fr/chemise-clement-kaa-coges-h12512.html> {'productname': 'Chemise Clément', 'gender': 'men', 'cleancategories': 'shirts', 'price': '180', 'color': 'VERT', 'picture_list': ['https://www.apc.fr/media/catalog/product/attribute/swatches_color/COGES_IAJ.jpg', 'https://www.apc.fr/media/catalog/product/attribute/swatches_color/COGES_KAA.jpg', 'https://www.apc.fr/media/catalog/product/cache/5f20f1917254e6a5a23af6773e8ed099/c/o/coges-h12512kaa_02_1684770825.jpg', 'https://www.apc.fr/media/catalog/product/cache/5f20f1917254e6a5a23af6773e8ed099/c/o/coges-h12512kaa_03_1684770825.jpg', 'https://www.apc.fr/media/catalog/product/cache/5f20f1917254e6a5a23af6773e8ed099/c/o/coges-h12512kaa_04_1684770825.jpg'], 'current_url': 'https://www.apc.fr/chemise-clement-kaa-coges-h12512.html'}
- Apify部署后异常:
parse函数正常执行(日志显示已解析列表页及商品页URL),但aaa_products函数从未触发,无对应日志输出,平台日志示例:
[apify] INFO TitleSpider is parsaaaaing <200 https://apify.com>... [apify] INFO TitleSpider is parsaaaaing <200 https://www.apc.fr/men/men-shirts.html>... [apify] INFO TitleSpider is parsaaaaing <200 https://www.apc.fr/chemise-clement-kaa-coevd-h12512.html>... [apify] INFO TitleSpider is parsaaaaing <200 https://www.apc.fr/surchemise-basile-pik-woapq-h02709.html>... [apify] INFO TitleSpider is parsaaaaing <200 https://www.apc.fr/chemise-greg-iaa-coguh-h12499.html>...
可能原因及修复方案
1. 商品链接未正确拼接(相对URL问题)
列表页提取的href可能是相对路径,本地环境中Scrapy会自动补全域名,但Apify环境下可能未处理,导致请求无效。
修复:用urljoin拼接完整URL:
# 在parse函数中替换原请求代码 from urllib.parse import urljoin full_url = urljoin(response.url, link) yield scrapy.Request( dont_filter=True, url=full_url, callback=self.aaa_products, )
2. 反爬拦截(请求头/IP识别)
Apify的默认请求头或IP可能被目标网站识别为爬虫,导致商品页请求返回空白页或验证页面(即使日志显示200状态码)。
验证与修复:
- 在
aaa_products开头添加日志,输出响应内容长度:
如果内容长度远小于正常页面,说明被拦截。可尝试:def aaa_products(self, response: Response): Actor.log.info(f'machin fait nimp {response}, content length: {len(response.text)}...') # 后续代码不变- 添加模拟浏览器的User-Agent请求头
- 在Apify中启用Headless Chrome替代纯Scrapy请求
- 使用Apify代理IP池
3. 代码缩进错误导致无输出
你的yield语句嵌套在for img_element in img_elements循环内部,若页面无img[data-src]元素,函数不会输出任何数据,且可能被视为无返回。
修复:将yield移到循环外部,确保即使没有图片也能输出商品基础信息:
def aaa_products(self, response: Response): Actor.log.info(f'machin fait nimp {response}...') # 原商品信息提取代码不变... picture_list = [] for img_element in img_elements: source_element = img_element.css('img::attr(data-src)').get() if source_element: # 增加非空判断 picture_list.append(source_element.strip()) # 将yield移到循环外部 yield { 'productname': productname, 'gender': gender, 'cleancategories': cleancategories, 'price': price, 'color': color, 'picture_list': picture_list, 'current_url': current_url, }
4. Referer头解析异常中断函数
你依赖Referer头提取gender和cleancategories,若Apify环境下请求未携带Referer,referer_url.decode('utf-8')会抛出AttributeError,导致函数中断。
修复:增加空值判断:
referer_url = response.request.headers.get('Referer', None) if not referer_url: Actor.log.warning(f'No Referer header for {response.url}') gender = None cleancategories = None else: referer_url = referer_url.decode('utf-8') split_url = referer_url.split("/") gender = split_url[-1].split("-")[0] last_part = split_url[-1] categorie = last_part.split("-")[1:] joinedcategories = "-".join(categorie) cleancategories = joinedcategories.replace(".html", "")
5. Scrapy与Apify异步环境冲突
导入的nest_asyncio和install_reactor可能与Apify的异步运行环境冲突,导致请求调度异常。
修复:调整启动逻辑,确保Scrapy在Apify环境中正确初始化,或移除不必要的 reactor 配置代码。
修复后完整爬虫代码
from typing import Generator from scrapy.responsetypes import Response from apify import Actor from urllib.parse import urljoin import nest_asyncio import scrapy from itemadapter import ItemAdapter from scrapy.crawler import CrawlerProcess from scrapy.utils.project import get_project_settings from scrapy.utils.reactor import install_reactor class TitleSpider(scrapy.Spider): name = 'title_spider' allowed_domains = ['apc.fr'] # 移到类属性,规范Scrapy配置 def start_requests(self): urls = [ "https://www.apc.fr/men/men-shirts.html" ] for url in urls: yield scrapy.Request( dont_filter=True, url=url, callback=self.parse, headers={ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } ) def parse(self, response: Response): Actor.log.info(f'TitleSpider is parsaaaaing {response}...') li_elements = response.css('li.product-item') for li in li_elements: productlink_container = li.css('.product-link') product_links = productlink_container.css('a::attr(href)').getall() for link in product_links: full_url = urljoin(response.url, link) yield scrapy.Request( dont_filter=True, url=full_url, callback=self.aaa_products, headers={ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': response.url } ) def aaa_products(self, response: Response): Actor.log.info(f'machin fait nimp {response}...') productname = response.css('h1.product-name::text').get() current_url = response.url productdescriptionfirst = response.css('div.product.attribute.intro') productdescriptionsecond = productdescriptionfirst.css('div.value') productdescription = productdescriptionsecond.css('::text').get() pricecontainer1 = response.css('div.product-price-wrapper') pricecontainer2 = pricecontainer1.css('div.price-final_price') pricecontainer3 = pricecontainer2.css('span.normal-price') pricecontainer4 = pricecontainer3.css('span[id^="product-price-"]') price = pricecontainer4.css('::attr(data-price-amount)').get() img_elements = response.css('img[data-src]') # 修复Referer解析逻辑 referer_url = response.request.headers.get('Referer', None) if not referer_url: Actor.log.warning(f'No Referer header for {response.url}') gender = None cleancategories = None else: referer_url = referer_url.decode('utf-8') split_url = referer_url.split("/") gender = split_url[-1].split("-")[0] last_part = split_url[-1] categorie = last_part.split("-")[1:] joinedcategories = "-".join(categorie) cleancategories = joinedcategories.replace(".html", "") colorcontainer3 = response.css('li.current-color-label') color = colorcontainer3.css('::text').get() picture_list = [] for img_element in img_elements: source_element = img_element.css('img::attr(data-src)').get() if source_element: picture_list.append(source_element.strip()) # 移到循环外yield yield { 'productname': productname, 'gender': gender, 'cleancategories': cleancategories, 'price': price, 'color': color, 'picture_list': picture_list, 'current_url': current_url, }
内容的提问来源于stack exchange,提问作者Thibault Bouveur
相关产品推荐
相关产品推荐

