如何在Python的BigQuery Storage API Writer中使用嵌套proto.Message?
使用proto-plus定义嵌套Protobuf消息写入BigQuery Storage的问题
使用proto-plus包在Python中定义Protobuf消息时,普通场景运行正常,但使用嵌套消息会出现问题。调用await bq_write_client.append_rows(iter([append_row_request]))时,会抛出以下错误:
google.api_core.exceptions.InvalidArgument: 400 Invalid proto schema: BqMessage.proto: Message.nested: "._default_package.Team" is not defined.
注:google-cloud-bigquery-storage库本身支持嵌套消息,官方示例使用独立的.proto文件,需要编译步骤,不如直接在Python中定义消息便捷。
初始代码示例
# Copyright 2021 Google LLC # # Licensed under the Apache License, Version 2.0 (the "License"); # you may not use this file except in compliance with the License. # You may obtain a copy of the License at # # https://www.apache.org/licenses/LICENSE-2.0 # # Unless required by applicable law or agreed to in writing, software # distributed under the License is distributed on an "AS IS" BASIS, # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. # See the License for the specific language governing permissions and # limitations under the License. import json import asyncio import proto from google.oauth2.service_account import Credentials from google.protobuf.descriptor_pb2 import DescriptorProto from google.cloud.bigquery_storage_v1beta2.types.storage import AppendRowsRequest from google.cloud.bigquery_storage_v1beta2.types.protobuf import ProtoSchema, ProtoRows from google.cloud.bigquery_storage_v1beta2.services.big_query_write import BigQueryWriteAsyncClient class Team(proto.Message): name = proto.Field(proto.STRING, number=1) class UserSchema(proto.Message): username = proto.Field(proto.STRING, number=1) email = proto.Field(proto.STRING, number=2) team = proto.Field(Team, number=3) async def main(): write_stream_path = BigQueryWriteAsyncClient.write_stream_path( "yolocommon", "test", "t_test_data", "_default") credentials = Credentials.from_service_account_file(filename="bigquery_config_file.json") bq_write_client = BigQueryWriteAsyncClient(credentials=credentials) proto_descriptor = DescriptorProto() UserSchema.pb().DESCRIPTOR.CopyToProto(proto_descriptor) proto_schema = ProtoSchema(proto_descriptor=proto_descriptor) serialized_rows = [] data = [ { "username": "Jack", "email": "jack@google.com", "nested": { "name": "Jack Jack" } }, { "username": "mary", "email": "mary@google.com", "nested": { "name": "Mary Mary" } } ] for item in data: instance = UserSchema.from_json(payload=json.dumps(item)) serialized_rows.append(UserSchema.serialize(instance)) proto_data = AppendRowsRequest.ProtoData( rows=ProtoRows(serialized_rows=serialized_rows), writer_schema=proto_schema ) append_row_request = AppendRowsRequest( write_stream=write_stream_path, proto_rows=proto_data ) result = await bq_write_client.append_rows(iter([append_row_request])) async for item in result: print(item) if __name__ == "__main__": asyncio.run(main())
更新1:调整嵌套消息定义
根据ProtoSchema文档要求:
输入消息的描述符必须是自包含的,仅使用单个描述符即可完全解码发送的数据行。对于由多个独立消息组成的数据行,这意味着描述符可能需要转换为仅使用嵌套类型。
调整消息定义为嵌套形式:
class UserSchema(proto.Message): class Team(proto.Message): name = proto.Field(proto.STRING, number=1) username = proto.Field(proto.STRING, number=1) email = proto.Field(proto.STRING, number=2) team = proto.Field(Team, number=3)
但仍抛出相同错误:
google.api_core.exceptions.InvalidArgument: 400 Invalid proto schema: BqMessage.proto: Message.nested: "._default_package.UserSchema.Team" is not defined.
更新2:问题根源
问题出在proto-plus的包名处理逻辑:当包名为空时,它会自动添加_default_package作为包名,这是为了规避基于upb的protobuf在添加到描述符池时的报错(错误信息为TypeError: Couldn't build proto file into descriptor pool: invalid name: empty part ()' means),相关代码中也有TODO注释说明,待protobuf修复后会恢复使用空字符串作为包值。
结论:目前无法使用proto.Message表示带有嵌套字段(STRUCT)的BigQuery表。
内容的提问来源于stack exchange,提问作者nodarai
相关产品推荐
相关产品推荐

