如何在Google Cloud Functions中用Flask+Scrapy实现爬虫服务
问题描述
需要在云端运行网页爬虫,接收携带邮编的POST请求,查询对应邮编并返回JSON格式的地址列表。目前仅拥有main.py文件和包含scrapy、flask的requirements.txt文件,但运行后出现多个错误,发送curl请求返回500状态码。
错误清单
TypeError: start_scrape() takes 0 positional arguments but 1 was givenMissingTargetException: File /workspace/main.py is expected to contain a function named /start_scrapeNameError: name 'start_urls' is not definedModuleNotFoundError: No module named 'scrapy'
修正后的完整代码
main.py
from multiprocessing import Process, Queue import scrapy from scrapy.crawler import CrawlerProcess from scrapy.utils.log import configure_logging import json # 独立定义爬虫类,便于复用和参数传递 class AddressesSpider(scrapy.Spider): name = 'Addresses' allowed_domains = ['find-energy-certificate.service.gov.uk'] def __init__(self, postcode, result_queue, *args, **kwargs): super().__init__(*args, **kwargs) self.start_urls = [f'https://find-energy-certificate.service.gov.uk/find-a-certificate/search-by-postcode?postcode={postcode}'] self.result_queue = result_queue self.addresses = [] def parse(self, response): # 跳过表格表头行 for row in response.xpath('//table[@class="govuk-table"]//tr[position()>1]'): address_text = row.xpath("normalize-space(.//a[@class='govuk-link']/text())").get() if not address_text: continue address = address_text.lower().rsplit(',', 2)[0] link = row.xpath('.//a[@class="govuk-link"]/@href').get() details = row.xpath("normalize-space(.//td/following-sibling::td)").get() self.addresses.append({ 'link': link, 'details': details, 'address': address }) def closed(self, reason): # 爬虫结束后将结果存入队列 self.result_queue.put(self.addresses) def run_spider(postcode): configure_logging() result_queue = Queue() def script(queue): try: process = CrawlerProcess(settings={ 'ROBOTSTXT_OBEY': False, 'LOG_LEVEL': 'ERROR' # 抑制冗余日志 }) # 向爬虫传入邮编和结果队列 process.crawl(AddressesSpider, postcode=postcode, result_queue=queue) process.start() except Exception as e: queue.put(e) main_process = Process(target=script, args=(result_queue,)) main_process.start() main_process.join() result = result_queue.get() if isinstance(result, Exception): raise result return result # 云端HTTP函数入口 def my_cloud_function(request): if request.method != 'POST': return ('Method not allowed', 405) try: request_data = request.get_json() postcode = request_data.get('postcode') if not postcode: return ('Missing postcode parameter', 400) addresses = run_spider(postcode) return (json.dumps(addresses), 200, {'Content-Type': 'application/json'}) except Exception as e: return (f'Error: {str(e)}', 500)
requirements.txt
scrapy>=2.8.0 flask>=2.3.0
错误修复详解
修复
start_scrape参数错误
Flask路由函数无需手动传入request参数,全局request对象可直接调用。同时云端HTTP函数(如GCP Cloud Function)只需通过指定入口函数my_cloud_function处理请求,无需额外Flask路由。修复云端函数入口名错误
云端函数入口必须是合法的函数名(不能带斜杠),部署时需在云平台配置中指定入口函数为my_cloud_function,而非/start_scrape。修复
start_urls未定义问题
将爬虫类独立于请求处理函数之外,通过构造函数传入postcode动态生成start_urls,并使用队列传递爬虫结果,避免变量作用域冲突。修复Scrapy依赖缺失问题
确保requirements.txt正确包含scrapy,云端平台会自动读取该文件安装依赖;本地测试时需先执行pip install -r requirements.txt安装依赖包。
部署与测试
云端部署(以GCP Cloud Function为例)
- 选择HTTP触发类型,设置入口函数为
my_cloud_function - 上传
main.py和requirements.txt,等待依赖安装完成
- 选择HTTP触发类型,设置入口函数为
测试请求
curl -H "Authorization: bearer $(gcloud auth print-identity-token)" \ -H "Content-Type: application/json" \ -d '{"postcode": "OX4+1EU"}' \ https://REGION-PROJECT_ID.cloudfunctions.net/FUNCTION_NAME替换
REGION、PROJECT_ID、FUNCTION_NAME为你的云函数实际信息。
内容的提问来源于stack exchange,提问作者Designer
相关产品推荐
相关产品推荐

