You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Google Cloud Functions中用Flask+Scrapy实现爬虫服务

问题描述

需要在云端运行网页爬虫,接收携带邮编的POST请求,查询对应邮编并返回JSON格式的地址列表。目前仅拥有main.py文件和包含scrapy、flask的requirements.txt文件,但运行后出现多个错误,发送curl请求返回500状态码。

错误清单

  • TypeError: start_scrape() takes 0 positional arguments but 1 was given
  • MissingTargetException: File /workspace/main.py is expected to contain a function named /start_scrape
  • NameError: name 'start_urls' is not defined
  • ModuleNotFoundError: No module named 'scrapy'

修正后的完整代码

main.py

from multiprocessing import Process, Queue
import scrapy
from scrapy.crawler import CrawlerProcess
from scrapy.utils.log import configure_logging
import json

# 独立定义爬虫类,便于复用和参数传递
class AddressesSpider(scrapy.Spider):
    name = 'Addresses'
    allowed_domains = ['find-energy-certificate.service.gov.uk']
    
    def __init__(self, postcode, result_queue, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.start_urls = [f'https://find-energy-certificate.service.gov.uk/find-a-certificate/search-by-postcode?postcode={postcode}']
        self.result_queue = result_queue
        self.addresses = []

    def parse(self, response):
        # 跳过表格表头行
        for row in response.xpath('//table[@class="govuk-table"]//tr[position()>1]'):
            address_text = row.xpath("normalize-space(.//a[@class='govuk-link']/text())").get()
            if not address_text:
                continue
            address = address_text.lower().rsplit(',', 2)[0]
            link = row.xpath('.//a[@class="govuk-link"]/@href').get()
            details = row.xpath("normalize-space(.//td/following-sibling::td)").get()

            self.addresses.append({
                'link': link,
                'details': details,
                'address': address
            })
    
    def closed(self, reason):
        # 爬虫结束后将结果存入队列
        self.result_queue.put(self.addresses)

def run_spider(postcode):
    configure_logging()
    result_queue = Queue()
    
    def script(queue):
        try:
            process = CrawlerProcess(settings={
                'ROBOTSTXT_OBEY': False,
                'LOG_LEVEL': 'ERROR'  # 抑制冗余日志
            })
            # 向爬虫传入邮编和结果队列
            process.crawl(AddressesSpider, postcode=postcode, result_queue=queue)
            process.start()
        except Exception as e:
            queue.put(e)
    
    main_process = Process(target=script, args=(result_queue,))
    main_process.start()
    main_process.join()
    
    result = result_queue.get()
    if isinstance(result, Exception):
        raise result
    return result

# 云端HTTP函数入口
def my_cloud_function(request):
    if request.method != 'POST':
        return ('Method not allowed', 405)
    
    try:
        request_data = request.get_json()
        postcode = request_data.get('postcode')
        if not postcode:
            return ('Missing postcode parameter', 400)
        
        addresses = run_spider(postcode)
        return (json.dumps(addresses), 200, {'Content-Type': 'application/json'})
    except Exception as e:
        return (f'Error: {str(e)}', 500)

requirements.txt

scrapy>=2.8.0
flask>=2.3.0

错误修复详解

  1. 修复start_scrape参数错误
    Flask路由函数无需手动传入request参数,全局request对象可直接调用。同时云端HTTP函数(如GCP Cloud Function)只需通过指定入口函数my_cloud_function处理请求,无需额外Flask路由。

  2. 修复云端函数入口名错误
    云端函数入口必须是合法的函数名(不能带斜杠),部署时需在云平台配置中指定入口函数为my_cloud_function,而非/start_scrape。

  3. 修复start_urls未定义问题
    将爬虫类独立于请求处理函数之外,通过构造函数传入postcode动态生成start_urls,并使用队列传递爬虫结果,避免变量作用域冲突。

  4. 修复Scrapy依赖缺失问题
    确保requirements.txt正确包含scrapy,云端平台会自动读取该文件安装依赖;本地测试时需先执行pip install -r requirements.txt安装依赖包。


部署与测试

  1. 云端部署(以GCP Cloud Function为例)

    • 选择HTTP触发类型,设置入口函数为my_cloud_function
    • 上传main.py和requirements.txt,等待依赖安装完成
  2. 测试请求

    curl -H "Authorization: bearer $(gcloud auth print-identity-token)" \
    -H "Content-Type: application/json" \
    -d '{"postcode": "OX4+1EU"}' \
    https://REGION-PROJECT_ID.cloudfunctions.net/FUNCTION_NAME
    

    替换REGION、PROJECT_ID、FUNCTION_NAME为你的云函数实际信息。

内容的提问来源于stack exchange,提问作者Designer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 12:45:34