You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫部署到Apify后无法执行aaa_products函数求助

Scrapy爬虫本地正常但部署到Apify后无法触发aaa_products函数的问题

问题现象

  • 本地运行正常:爬虫可正常解析商品页面,输出预期结果,日志示例:
2023-10-08 00:15:41 [scrapy.core.scraper] DEBUG: Scraped from <200 https://www.apc.fr/chemise-clement-kaa-coges-h12512.html>
{'productname': 'Chemise Clément', 'gender': 'men', 'cleancategories': 'shirts', 'price': '180', 'color': 'VERT', 'picture_list': ['https://www.apc.fr/media/catalog/product/attribute/swatches_color/COGES_IAJ.jpg', 'https://www.apc.fr/media/catalog/product/attribute/swatches_color/COGES_KAA.jpg', 'https://www.apc.fr/media/catalog/product/cache/5f20f1917254e6a5a23af6773e8ed099/c/o/coges-h12512kaa_02_1684770825.jpg', 'https://www.apc.fr/media/catalog/product/cache/5f20f1917254e6a5a23af6773e8ed099/c/o/coges-h12512kaa_03_1684770825.jpg', 'https://www.apc.fr/media/catalog/product/cache/5f20f1917254e6a5a23af6773e8ed099/c/o/coges-h12512kaa_04_1684770825.jpg'], 'current_url': 'https://www.apc.fr/chemise-clement-kaa-coges-h12512.html'}
  • Apify部署后异常:parse函数正常执行(日志显示已解析列表页及商品页URL),但aaa_products函数从未触发,无对应日志输出,平台日志示例:
[apify] INFO  TitleSpider is parsaaaaing <200 https://apify.com>...
[apify] INFO  TitleSpider is parsaaaaing <200 https://www.apc.fr/men/men-shirts.html>...
[apify] INFO  TitleSpider is parsaaaaing <200 https://www.apc.fr/chemise-clement-kaa-coevd-h12512.html>...
[apify] INFO  TitleSpider is parsaaaaing <200 https://www.apc.fr/surchemise-basile-pik-woapq-h02709.html>...
[apify] INFO  TitleSpider is parsaaaaing <200 https://www.apc.fr/chemise-greg-iaa-coguh-h12499.html>...

可能原因及修复方案

1. 商品链接未正确拼接(相对URL问题)

列表页提取的href可能是相对路径,本地环境中Scrapy会自动补全域名,但Apify环境下可能未处理,导致请求无效。

修复:用urljoin拼接完整URL:

# 在parse函数中替换原请求代码
from urllib.parse import urljoin

full_url = urljoin(response.url, link)
yield scrapy.Request(
    dont_filter=True, 
    url=full_url, 
    callback=self.aaa_products,
)

2. 反爬拦截(请求头/IP识别)

Apify的默认请求头或IP可能被目标网站识别为爬虫,导致商品页请求返回空白页或验证页面(即使日志显示200状态码)。

验证与修复:

  • 在aaa_products开头添加日志,输出响应内容长度:
    def aaa_products(self, response: Response):
        Actor.log.info(f'machin fait nimp {response}, content length: {len(response.text)}...')
        # 后续代码不变
    
    如果内容长度远小于正常页面,说明被拦截。可尝试:
    • 添加模拟浏览器的User-Agent请求头
    • 在Apify中启用Headless Chrome替代纯Scrapy请求
    • 使用Apify代理IP池

3. 代码缩进错误导致无输出

你的yield语句嵌套在for img_element in img_elements循环内部,若页面无img[data-src]元素,函数不会输出任何数据,且可能被视为无返回。

修复:将yield移到循环外部,确保即使没有图片也能输出商品基础信息:

def aaa_products(self, response: Response):
    Actor.log.info(f'machin fait nimp {response}...')
    
    # 原商品信息提取代码不变...
    
    picture_list = []
    for img_element in img_elements:
        source_element = img_element.css('img::attr(data-src)').get()
        if source_element:  # 增加非空判断
            picture_list.append(source_element.strip())          
    
    # 将yield移到循环外部
    yield {
        'productname': productname,
        'gender': gender,
        'cleancategories': cleancategories,
        'price': price,
        'color': color,
        'picture_list': picture_list,
        'current_url': current_url,
    }

4. Referer头解析异常中断函数

你依赖Referer头提取gender和cleancategories,若Apify环境下请求未携带Referer,referer_url.decode('utf-8')会抛出AttributeError,导致函数中断。

修复:增加空值判断:

referer_url = response.request.headers.get('Referer', None)
if not referer_url:
    Actor.log.warning(f'No Referer header for {response.url}')
    gender = None
    cleancategories = None
else:
    referer_url = referer_url.decode('utf-8')
    split_url = referer_url.split("/")
    gender = split_url[-1].split("-")[0]
    last_part = split_url[-1]
    categorie = last_part.split("-")[1:]
    joinedcategories = "-".join(categorie)
    cleancategories = joinedcategories.replace(".html", "")

5. Scrapy与Apify异步环境冲突

导入的nest_asyncio和install_reactor可能与Apify的异步运行环境冲突,导致请求调度异常。

修复:调整启动逻辑,确保Scrapy在Apify环境中正确初始化,或移除不必要的 reactor 配置代码。

修复后完整爬虫代码

from typing import Generator
from scrapy.responsetypes import Response
from apify import Actor
from urllib.parse import urljoin
import nest_asyncio
import scrapy
from itemadapter import ItemAdapter
from scrapy.crawler import CrawlerProcess
from scrapy.utils.project import get_project_settings
from scrapy.utils.reactor import install_reactor


class TitleSpider(scrapy.Spider):
    name = 'title_spider'
    allowed_domains = ['apc.fr']  # 移到类属性,规范Scrapy配置

    def start_requests(self):
        urls = [
            "https://www.apc.fr/men/men-shirts.html"
        ]
        for url in urls:
            yield scrapy.Request(
                dont_filter=True, 
                url=url, 
                callback=self.parse,
                headers={
                    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
                }
            )

    def parse(self, response: Response):
        Actor.log.info(f'TitleSpider is parsaaaaing {response}...')
        li_elements = response.css('li.product-item')

        for li in li_elements:
            productlink_container = li.css('.product-link')
            product_links = productlink_container.css('a::attr(href)').getall()

            for link in product_links:
                full_url = urljoin(response.url, link)
                yield scrapy.Request(
                    dont_filter=True, 
                    url=full_url, 
                    callback=self.aaa_products,
                    headers={
                        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
                        'Referer': response.url
                    }
                )

    def aaa_products(self, response: Response):
        Actor.log.info(f'machin fait nimp {response}...')
    
        productname = response.css('h1.product-name::text').get()
        current_url = response.url
        productdescriptionfirst = response.css('div.product.attribute.intro')
        productdescriptionsecond = productdescriptionfirst.css('div.value')
        productdescription = productdescriptionsecond.css('::text').get()
        pricecontainer1 = response.css('div.product-price-wrapper')
        pricecontainer2 = pricecontainer1.css('div.price-final_price')
        pricecontainer3 = pricecontainer2.css('span.normal-price')
        pricecontainer4 = pricecontainer3.css('span[id^="product-price-"]')
        price = pricecontainer4.css('::attr(data-price-amount)').get()
        img_elements = response.css('img[data-src]')
        
        # 修复Referer解析逻辑
        referer_url = response.request.headers.get('Referer', None)
        if not referer_url:
            Actor.log.warning(f'No Referer header for {response.url}')
            gender = None
            cleancategories = None
        else:
            referer_url = referer_url.decode('utf-8')
            split_url = referer_url.split("/")
            gender = split_url[-1].split("-")[0]
            last_part = split_url[-1]
            categorie = last_part.split("-")[1:]
            joinedcategories = "-".join(categorie)
            cleancategories = joinedcategories.replace(".html", "")
        
        colorcontainer3 = response.css('li.current-color-label')
        color = colorcontainer3.css('::text').get()
        
        picture_list = []
        for img_element in img_elements:
            source_element = img_element.css('img::attr(data-src)').get()
            if source_element:
                picture_list.append(source_element.strip())          
        
        # 移到循环外yield
        yield {
            'productname': productname,
            'gender': gender,
            'cleancategories': cleancategories,
            'price': price,
            'color': color,
            'picture_list': picture_list,
            'current_url': current_url,
        }

内容的提问来源于stack exchange,提问作者Thibault Bouveur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 00:05:55