GCP云函数URL状态检测调用返回403错误问题咨询
403报错根因排查
- 你当前给服务账号绑定的
Cloud Functions Viewer属于云函数资源管控类权限,仅支持查看云函数的配置、元数据信息,不具备调用云函数HTTP触发端点的权限,无法通过调用鉴权校验。 - 你虽然在环境变量中配置了服务账号密钥路径,但
requests库默认不会自动加载Google身份凭证、在请求中添加合法的Authorization鉴权头,相当于发起的是匿名请求,默认配置(需认证)的云函数会直接返回403。 - 额外排查项:如果云函数配置了入站流量限制(如仅允许VPC内部访问、组织策略禁止公网暴露),公网发起的请求同样会返回403类错误。
- 若你需要公网匿名访问云函数,可在云函数权限页给
allUsers主体绑定Cloud Functions Invoker角色,但生产环境不推荐该配置,存在未授权访问风险。
整体实现方案
你的需求(检测云函数可用性、接入GCP监控展示)完全可落地,整体流程为:带鉴权探测云函数端点 -> 采集状态码、响应延迟、可用性指标 -> 写入GCP Cloud Monitoring -> 配置可视化看板与告警。
1. 权限前置配置
- 给执行探测逻辑的服务账号,针对目标云函数绑定
Cloud Functions Invoker角色,遵循最小权限原则不要过度授权。 - 确认云函数入站规则允许探测请求的源地址访问:探测脚本跑在公网则开启公网访问权限,跑在GCP内部资源则配置对应VPC访问规则。
- 给该服务账号额外绑定
Monitoring Metric Writer角色,用于向Cloud Monitoring写入自定义监控指标。
2. 参考实现代码
首先安装依赖:
pip install requests google-auth google-cloud-monitoring
探测+指标上报代码:
import os import time import requests from google.auth.transport.requests import Request as AuthRequest from google.oauth2 import id_token from google.cloud import monitoring_v3 # 基础配置 os.environ["GOOGLE_APPLICATION_CREDENTIALS"] = "sa.json" CF_REGION = "替换为目标云函数所在区域,例如us-central1" CF_PROJECT_ID = "替换为目标云函数所属项目ID" CF_NAME = "替换为目标云函数名称" METRIC_PROJECT_ID = "替换为存储监控指标的项目ID,通常与云函数项目ID一致" CF_URL = f"https://{CF_REGION}-{CF_PROJECT_ID}.cloudfunctions.net/{CF_NAME}" # 生成云函数调用要求的OIDC鉴权头 def gen_auth_headers(): auth_req = AuthRequest() # OIDC Token的audience必须配置为目标云函数的URL token = id_token.fetch_id_token(auth_req, CF_URL) return {"Authorization": f"Bearer {token}"} # 向Cloud Monitoring写入自定义指标 def upload_metrics(status_code: int, latency_ms: float, is_up: int): client = monitoring_v3.MetricServiceClient() project_path = f"projects/{METRIC_PROJECT_ID}" now_ts = time.time() ts_sec = int(now_ts) ts_nano = int((now_ts - ts_sec) * 10**9) # 三个核心指标:响应状态码、响应延迟、可用状态(1=正常 0=异常) metric_list = [ ("custom.googleapis.com/cf_healthcheck/status", status_code), ("custom.googleapis.com/cf_healthcheck/latency_ms", latency_ms), ("custom.googleapis.com/cf_healthcheck/up", is_up) ] for metric_type, val in metric_list: time_series = monitoring_v3.TimeSeries() time_series.metric.type = metric_type time_series.resource.type = "generic_task" time_series.resource.labels["project_id"] = METRIC_PROJECT_ID time_series.resource.labels["location"] = CF_REGION time_series.resource.labels["namespace"] = "cf_probe" time_series.resource.labels["job"] = CF_NAME time_series.resource.labels["task_id"] = "healthcheck_task" point = monitoring_v3.Point() point.value.double_value = val point.interval.end_time.seconds = ts_sec point.interval.end_time.nanos = ts_nano time_series.points.append(point) client.create_time_series(name=project_path, time_series=[time_series]) # 单次健康探测逻辑 def run_probe(): headers = gen_auth_headers() start = time.time() up_status = 0 resp_code = 0 try: resp = requests.get(CF_URL, headers=headers, timeout=10) cost_ms = (time.time() - start) * 1000 resp_code = resp.status_code # 可用判定规则:2xx、3xx状态码判定为正常,可按需调整规则 if 200 <= resp_code < 400: up_status = 1 except Exception as e: cost_ms = (time.time() - start) * 1000 print(f"Probe request failed: {str(e)}") upload_metrics(resp_code, cost_ms, up_status) print(f"Probe result: code={resp_code}, cost={cost_ms:.2f}ms, available={up_status}") if __name__ == "__main__": run_probe()
3. 监控配置指引
- 将探测逻辑部署为定时任务:可搭配GCP Cloud Scheduler,按固定间隔(如1分钟)触发一个轻量Cloud Function/Cloud Run Job执行探测代码,不需要长期运行服务器。
- 可视化看板:进入GCP Cloud Monitoring控制台,使用上报的三个自定义指标
cf_healthcheck/up、cf_healthcheck/latency_ms、cf_healthcheck/status配置折线图、可用性百分比统计卡片。 - 告警配置:可创建告警策略,当连续2次探测
up指标值为0时,通过邮件、短信、企业IM渠道发送告警通知。
内容的提问来源于stack exchange,提问作者ash_ketchum12
相关产品推荐
相关产品推荐

