BigQuery Python API执行JSON Schema文件创建表失败问题求助
解决BigQuery Python API创建表时的Schema格式错误问题
我一眼就看出问题出在你读取schema文件的方式上——你用readlines()把JSON文件读成了字符串列表,但BigQuery Python API的Table构造器根本不认这个格式,它需要的是解析后的Python字典列表,或者是bigquery.SchemaField对象列表。
问题根源
fName.readlines()返回的是文件每行的原始字符串,比如类似这样的内容:
['[', ' { "name":"user_id", "type":"STRING", "mode":"NULLABLE" },', ...]
这些字符串对BigQuery API来说就是一堆无意义的文本,自然会抛出Schema items must either be fields or compatible mapping representations的错误。而控制台和CLI能正常使用这个文件,是因为它们内部会自动帮你完成JSON到结构化数据的解析。
修复方案
你只需要导入json模块,用json.load()来加载并解析整个JSON文件,把它转换成Python能识别的列表和字典结构。修改后的代码如下:
import json from google.cloud import bigquery # Construct a BigQuery client object. client = bigquery.Client() table_id = "project-py-290522:bq_dts.bq-test" def open_schema(): with open("hcl-schema.json","r", encoding = "utf-8") as fName: schema = json.load(fName) # 替换readlines为json.load table = bigquery.Table(table_id, schema=schema) print(repr(table)) client.create_table(table) # Make an API request. return table # 返回table对象,让外部的print能访问到 if __name__ == "__main__": table = open_schema() print("Created table {}.{}.{}".format(table.project, table.dataset_id, table.table_id))
额外说明
另外注意你原来代码里的一个小问题:print("Created table...")放在了if __name__ == "__main__"块外面,这样即使不执行主逻辑也会打印,而且还会因为table变量未定义报错。我已经把它移到了块内,并且让open_schema()返回table对象,这样就能正常输出创建成功的信息了。
这样修改后,你的脚本就能像控制台和CLI一样,正确识别schema并创建表了。
内容的提问来源于stack exchange,提问作者Dean Geary
相关产品推荐
相关产品推荐

