Google Cloud Build部署Julia应用至Cloud Run访问报500错误
问题描述
我在Google Cloud Build上部署Julia应用,已通过现有配置文件完成Cloud Build镜像构建,且成功部署到Cloud Run服务,但访问Cloud Run分配的托管地址https://backend.app时返回500错误,错误截图如下:
错误截图
项目全部配置文件如下:
Dockerfile
FROM julia:1.6.5 RUN apt-get update && apt-get install -y gcc ENV JULIA_PROJECT @. WORKDIR /home ENV VERSION 1 ADD . /home RUN julia deploy/packagecompile.jl EXPOSE 9086 ENTRYPOINT ["julia", "-JApp.so", "-t", "auto", "-L", "src/backend.jl", "-e", "Backend.startapp()"]
cloudbuild.yaml
steps: - name: gcr.io/cloud-builders/docker args: - build - '-t' - 'gcr.io/project-id/github.com/username/appname:$COMMIT_SHA' - ./backend - name: gcr.io/cloud-builders/docker args: - push - 'gcr.io/project-id/github.com/username/appname:$COMMIT_SHA' - name: gcr.io/cloud-builders/gcloud args: - run - deploy - backend - '--image' - 'gcr.io/project-id/github.com/username/appname:$COMMIT_SHA' - '--region' - us-global - '--platform' - managed timeout: 3600s images: - gcr.io/project-id/github.com/username/appname
cloudrun.yaml
apiVersion: serving.knative.dev/v1 kind: Service metadata: name: backend namespace: 'xxxxxxxx' selfLink: /apis/serving.knative.dev/v1/namespaces/xxxxxx/services/backend uid: xxxxxxxxxxxx-xxxxxxxxxxxxxxx-xxxxxxxxxxxxxx-xxxxxxx-xxxxxxxxxxxx resourceVersion: xxxxxxx generation: 8 creationTimestamp: '2022-06-13T02:26:09.088993Z' labels: managed-by: gcp-cloud-build-deploy-cloud-run gcb-trigger-id: xxxxxxxxxxxxxxxxxx-xxxxxxxxxxxxxx-xxxxxxxxxx-xxxxxxxxx-xxxxxxxxxxxxx cloud.googleapis.com/location: us-global annotations: run.googleapis.com/client-name: gcloud serving.knative.dev/creator: xxxxxxxxxxx serving.knative.dev/lastModifier: xxxxxx@cloudbuild.gserviceaccount.com client.knative.dev/user-image: gcr.io/project-id/github.com/username/appname:xxxxxxxxxxxxxx run.googleapis.com/client-version: 387.0.0 run.googleapis.com/ingress: all run.googleapis.com/ingress-status: all spec: template: metadata: name: app-backend-0032348-dob annotations: run.googleapis.com/client-name: gcloud client.knative.dev/user-image: gcr.io/project-id/github.com/username/appname:xxxxxxxxxxxxxx run.googleapis.com/client-version: 387.0.0 autoscaling.knative.dev/minScale: '1' autoscaling.knative.dev/maxScale: '100' spec: containerConcurrency: 80 timeoutSeconds: 300 serviceAccountName: xxxxxxxxxxxx containers: - image: gcr.io/project-id/github.com/username/appname:xxxxxxxxxxxxxx ports: - name: http1 containerPort: 9086 resources: limits: cpu: 1000m memory: 512Mi traffic: - percent: 100 latestRevision: true status: observedGeneration: 8 conditions: - type: Ready status: 'True' lastTransitionTime: '2022-06-13T09:02:43.268372Z' - type: ConfigurationsReady status: 'True' lastTransitionTime: '2022-06-13T09:02:36.523473Z' - type: RoutesReady status: 'True' lastTransitionTime: '2022-06-13T09:02:43.268372Z' latestReadyRevisionName: app-backend-0032348-dob latestCreatedRevisionName: app-backend-0032348-dob traffic: - revisionName: app-backend-0032348-dob percent: 100 latestRevision: true url: https://backend.app address: url: https://backend.app
问题排查过程
首次日志排查
获取Cloud Run服务端日志,开头的沙箱不支持系统调用提示为无害信息可直接忽略,核心错误为Nettle_jll初始化失败,提示Nettle制品未正确安装:
Container Sandbox: Unsupported syscall UNKNOWN[1008/0x3f0](0x0,0x0,0x0,0x0,0x0,0x0). It is very likely that you can safely ignore this message and that this is not the cause of any error you might be troubleshooting. fatal: error thrown and no exception handler available. InitError(mod=:Nettle_jll, error=ErrorException("Artifact "Nettle" was not installed correctly. Try `using Pkg; Pkg.instantiate()` to re-install all missing resources.")) error at ./error.jl:33 # 省略部分调用堆栈 Container called exit(1).
构建过程排查
使用strace跟踪构建流程,构建阶段末尾出现超时错误:
# 省略部分系统调用日志 Step #0 - "Build": [0m[91m[pid 320] brk(0x2a6c8000) = 0x2a6c8000 TIMEOUT ERROR: context deadline exceeded Step #0 - "Build": [0m
初步推测为PackageCompiler无法找到匹配的CPU目标导致构建失败。后续在本地构建镜像推送至GCP部署时,出现同类错误:
ERROR: Unable to find compatible target in system image. Uncaught signal: 11, pid=1, tid=1, fault_addr=6920. Container terminated on signal 11.
尝试显式传入cpu_target参数修复该问题,但该方案仅在Linux环境下构建成功,Windows环境下构建直接失败。
可行解决方案
- 修复依赖制品缺失问题:在Dockerfile的包编译执行步骤前,新增命令
RUN julia -e 'using Pkg; Pkg.instantiate(); Pkg.precompile()',保证所有依赖、JLL二进制制品在镜像构建阶段就完整安装,避免运行时查找制品路径失败。 - 解决CPU指令集不兼容问题:调用PackageCompiler生成系统镜像时,显式传入参数
cpu_target="generic",生成兼容所有x86_64架构的通用系统镜像,既可以适配Cloud Run底层gvisor沙箱的CPU指令集,也支持Windows、macOS、Linux多环境本地构建,不会出现跨环境构建失败或运行时找不到兼容目标的问题。 - 调整构建超时配置:当前Cloud Build设置的3600s超时容易在Julia包编译阶段触发中断,将超时时间调整为7200s即可,若编译过程出现内存不足,可同步提升Cloud Build构建机的内存配置。
- 校验服务监听配置:当前配置的容器暴露端口9086和Cloud Run转发端口一致无需修改,但需要确认
Backend.startapp()启动的HTTP服务监听0.0.0.0地址,而非本地回环地址127.0.0.1,否则Cloud Run侧无法正常转发请求,也会返回500错误。
内容的提问来源于stack exchange,提问作者Mohammad Saad
相关产品推荐
相关产品推荐

