You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

经Cloud Function-Protobuf-PubSub流转后BigQuery存0.0/false为null的原因

问题:Protobuf编码后BigQuery中默认值字段显示为null

我使用Python 3.10运行时的Cloud Function接收JSON payload,按照指定Protobuf Schema编码后发布至PubSub Topic,最终将数据处理至BigQuery。但处理完成后,John的is_active字段和Sam的fee字段显示为null,而非预期的false和0.0。

原始JSON Payload

{
  "data": [
    {
      "user_id": "XY25999A",
      "firstname": "John",
      "lastname": "Doe",
      "fee": 20.00,
      "is_active": false
    },
    {
      "user_id": "XY26999B",
      "firstname": "Sam",
      "lastname": "Foo",
      "fee": 0.00,
      "is_active": true
    },
    {
      "user_id": "XY27999C",
      "firstname": "Kay",
      "lastname": "Bent",
      "fee": 20.00,
      "is_active": true
    }
  ]
}

JSON Schema

{
    "type":"object",
    "properties":{
       "user_id":{
          "type":"string"
       },
       "firstname":{
          "type":"string"
       },
       "lastname":{
          "type":"string"
       },
       "fee":{
          "type":"number"
       },
       "is_active":{
          "type":"boolean"
       }
    }
}

Protobuf Schema

message ProtoSchema {
    string user_id = 1;
    string firstname = 2;
    string lastname = 3;
    double fee = 4;
    bool is_active = 5;
}

BigQuery处理结果

user_idfirstnamelastnamefeeis_active
XY25999AJohnDoe20.00null
XY26999BSamFoonulltrue
XY27999CKayBent20.00true

原因分析

这个问题的核心是Protobuf的默认值序列化规则:

  • Protobuf对基础类型预设了默认值:bool类型默认值为false,double类型默认值为0.0。
  • 当字段值等于对应类型的默认值时,Protobuf在序列化过程中会省略该字段,不会将其写入最终的二进制数据。

对应到你的场景:

  1. John的is_active值为false,正好匹配bool类型的默认值,因此该字段未被Protobuf序列化,PubSub传递的消息中不包含这个字段,BigQuery解析时会将缺失字段填充为null。
  2. Sam的fee值为0.0,正好匹配double类型的默认值,同样未被序列化,导致BigQuery中该字段显示为null。

内容的提问来源于stack exchange,提问作者Banty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 07:00:58